Blog

2026.08.17

Tech

Introducing Matlantis-PFP v9: Benchmarking on MLIP Arena and Improving Experimental Agreement with r2SCAN

Nontawat Charoenphakdee

Overview

Since the release of Matlantis in July 2021, Matlantis-PFP [1] (hereafter referred to as PFP), the core Machine Learning Interatomic Potential (MLIP) technology behind Matlantis™, has continued to evolve through updates. Each release has expanded the applicability of PFP and improved its accuracy across a wide range of atomistic simulation tasks. In this article, we introduce the latest release, PFP v9.0.0, highlighting its key updates and briefly positioning its performance using the score from MLIP Arena [2], a public benchmark for universal MLIPs (uMLIPs).

Key Improvements in PFP v9.0.0

In recent years, universal MLIPs have demonstrated their high expressive capabilities in practice, accelerating their application to real-world problems. This process has revealed numerous new challenges related to real-world applicability, and these aspects are increasingly recognized as important performance criteria for MLIPs. As a result, benchmark development and MLIP advancement are now being driven by this focus on practical applications.

Recently, we announced the release of PFP v9.0.0. This represents our latest iteration of the uMLIP framework, which has been extensively utilized across diverse real-world tasks through Matlantis since 2021.

Building on PFP v8.0.0, which introduced an r2SCAN-functional-based mode to improve agreement with experimental values or high-accuracy references [3], PFP v9.0.0 further expands the elemental coverage of this mode by adding lanthanide- and actinide-series elements. As a result, the total number of supported elements has increased from 70 to 96, ranging from hydrogen (H) to curium (Cm). This broad elemental coverage enables accurate simulations across a wide range of materials, from common organic and inorganic systems to materials containing lanthanides and actinides.

Our latest training data for PFP v9.0.0 includes not only crystals and molecules, but also diverse atomic structures such as surface slabs, adsorption structures, clusters, and metal complexes. This diversity is important because practical materials simulations often involve environments that differ substantially from ideal bulk crystals. Surfaces, interfaces, adsorbates, and nanoclusters frequently appear in catalysis, battery materials, alloys, porous materials, and other application areas. By incorporating these structures into the training data, PFP v9.0.0 is designed to provide reliable predictions across a broader range of chemical and structural situations.

Compared with PFP v8.0.0, PFP v9.0.0 shows improvements in the GMTKN55 molecular benchmark [3], melting points, and surface energies, while maintaining comparable performance for formation energies and phase diagrams.

MLIP Arena benchmark results

While the key improvements highlight the r2SCAN performance, the lack of a public r2SCAN-based benchmark makes it difficult to place these gains in a broader context. Therefore, to demonstrate the capability of PFP v9 on a public benchmark, we evaluate it on the well-known MLIP Arena proposed by Chiang et al. [2], specifically using the PBE mode of PFP v9. MLIP Arena is an open benchmarking platform designed to assess machine-learning interatomic potentials beyond conventional static error metrics, such as energy and force prediction errors. Instead, it evaluates whether a model reproduces physically meaningful behavior in practical simulation settings, providing a complementary measure of model reliability for downstream atomistic simulations.

Following the MLIP Arena protocol, we compare model performance across five benchmark tasks: homonuclear diatomics (diatomics), equation of state (eos_bulk), energy-volume scans (wbm_ev), stability, and combustion. For detailed definitions of each task, please refer to the original MLIP Arena paper [2].

Table 1: Overall ranking

PFP v9 achieved the 1st-place result, tied with MACE-MPA in the overall MLIP Arena ranking. The leaderboard values for models other than PFP v9.0.0 are based on the public MLIP Arena leaderboard as of March 1, 2026. This result places PFP v9 at the top of the arena, among the strongest models evaluated under the MLIP Arena protocol.

Table 2: Head-to-head comparison between PFP and MACE-MPA

Since overall rankings may shift as new models enter the arena, we also examine head-to-head comparisons (i.e., pairwise comparisons between any two models). Notably, PFP v9 ranks first in three of the five tasks, guaranteeing it a majority victory in every pairwise comparison against any model in the arena, including MACE-MPA: an opponent can win at most the remaining two tasks. Consequently, no model in the arena can surpass PFP v9 in head-to-head evaluation. Detailed task-level results are available in this blog post.

Future development direction

As uMLIPs move toward practical use, the range of capabilities that need to be evaluated is also expanding. Traditional benchmarks based on regression accuracy for energies and forces have been essential for demonstrating the expressive power of uMLIPs. In real materials research, however, models must also be evaluated from more application-oriented perspectives, including stability in long-time simulations, properties obtained from long-horizon tasks such as MD simulations, and agreement with macroscopic experimental values. We believe that this deepening of benchmarks is similar to what has been observed in the development of other foundation models, including Large Language Models (LLMs), and reflects the fact that uMLIPs are moving toward a more practical stage.

In this context, next-generation benchmarks for uMLIPs have two important roles. One is to compare different models and provide insights for model development. The other, which is especially important for users, is to clarify where current uMLIPs work well and where caution is still needed. Such information helps users apply uMLIPs with greater confidence and helps developers identify the next challenges. MLIP Arena is a pioneering effort in this direction, proposing metrics that place importance on simulation stability and practical behavior. These metrics are well aligned with the direction we are pursuing. At the same time, practical-task-oriented benchmarking must cover a wide range of application cases, and many challenges remain.

With this perspective, we are currently developing an internal benchmark environment based on the following criteria:

  • Evaluation of practical task behavior, including long-time simulations, rather than only single-point calculations
  • Evaluation of properties that affect actual phenomena, such as reaction energies and defect formation energies
  • Metrics that avoid excessive optimization to a particular DFT functional, with attention to experimental values
  • Coverage of important simulation cases across diverse materials
  • Evaluation of simulation stability and unexpected anomalous behavior
  • Active inclusion of cases that current uMLIPs cannot yet evaluate accurately

These benchmarks are currently being developed mainly for internal evaluation. In our recent preprint on PFP v8 [3], we reported several results from practical calculation tasks. Because evaluations close to real applications are also important for the broader community, we plan to release what we can as preparations are completed. By continuing this cycle of application-oriented evaluation and improvement, we aim to develop Matlantis-PFP into a foundational technology that materials researchers can use with greater confidence in real materials development.

References:

[1] S. Takamoto et al., Nat. Commun. 13, 2991 (2022). https://www.nature.com/articles/s41467-022-30687-9
[2] Y. Chiang et al., NeurIPS 38 (2025). https://papers.nips.cc/paper_files/paper/2025/hash/bfa45223cc236855dbaa5c468c809896-Abstract-Datasets_and_Benchmarks_Track.html (https://huggingface.co/spaces/atomind/mlip-arena)
[3] C. Shinagawa et al., “Matlantis-PFP v8: Universal Machine Learning Interatomic Potential with Better Experimental Agreements via r2SCAN Functional”, arXiv:2603.11063 (2026). https://arxiv.org/abs/2603.11063

  • Twitter
  • Facebook