MatFlowBench evaluates how reliably and efficiently LLM agents carry out the materials science workflows of computational and experimental scientists. Agents run multi-step calculations, request experimental measurements, and interpret the results to reach scientific conclusions.
02 / The constructWhat a successful task establishes
Completing a research workflow.
MatFlowBench measures whether an LLM agent can complete a specified materials research task using the available tools and data, within the task’s resource limits.
Each task’s executable verifier compares results with numerical references, using tolerances set by the underlying physics. Other checks cover physical constraints from well-established scientific workflows. Additionally, some tasks require calculation files or measurement logs.
A pass means that the submitted result meets these checks. It does not by itself establish new materials discovery, complete physical validation, or sound reasoning at every step.
01 Interpret the objective02 Use tools and data03 Submit the result04 Check the result
03 / The workflowsTwo evidence settings
Computational and experimental workflows.
In silico25 tasks
Computational
Agents are given simulation software and a scientific goal. They plan and run the calculations themselves, much as a computational materials scientist would and must carry the work through to the property or conclusion the task asks for.
From measured data20 tasks
Experimental
Agents receive a set of target properties and access to a simulated lab, where every experiment returns real measured data. Each sample can be characterized on several instruments, and each instrument reveals a different set of properties. With a finite budget to spend, the agent decides which samples to prepare and which measurements are worth running, using what it learns to narrow the search until it finds a material that meets the requirements.
04 / EvaluationHow a rollout is scored
What the score means.
Each rollout passes or fails. The task’s executable verifier makes that call, and the leaderboard reports how often each model passes.
01
Numerical references
Results are compared with numerical references, using tolerances set by the underlying physics.
02
Physical constraints and evidence
Other checks cover physical constraints from well-established scientific workflows. Some tasks also require calculation files or measurement logs, and experimental tasks must reach a qualifying material within the measurement budget.
03
Pass rate and Pass@1
Each model ran with high reasoning effort, twice per computational task and four times per experimental task. Pass rate is the share of all rollouts that passed, and the leaderboard is ranked by it. Pass@1 is the share of passing rollouts on a task, averaged so that every task counts equally.