Materials WorkFlow Benchmark

A benchmark for
materials research.

MatFlowBench evaluates how reliably and efficiently LLM agents carry out the materials science workflows of computational and experimental scientists. Agents run multi-step calculations, request experimental measurements, and interpret the results to reach scientific conclusions.

45 candidate tasks · Private benchmark

Model performance.

MatFlowBench leaderboard: Total
#ModelHarnessReasoningPass rateRollouts passedPass@1Tasks solved
1GPT-6 AstraOpenAICodexHigh60.8%79 / 13060.0%34 / 45
2Claude Opus 5.5AnthropicClaude CodeHigh53.8%70 / 13056.7%30 / 45
3Claude Sonnet 5.5AnthropicClaude CodeHigh47.7%62 / 13050.0%32 / 45
4GPT-6.1 SolOpenAICodexHigh44.6%58 / 13045.0%30 / 45
5Grok 4.7xAITerminus 2High26.2%34 / 13022.8%17 / 45
6Gemini 3.8 FlashGoogleTerminus 2High25.4%33 / 13018.9%14 / 45
7Kimi K3Moonshot AITerminus 2High23.8%31 / 13022.2%16 / 45
8Tencent HY4 PreviewTencentTerminus 2High22.3%29 / 13020.0%16 / 45
9Muse Spark 1.3MetaTerminus 2High20.0%26 / 13016.7%15 / 45
10Qwen3.8-27BAlibabaTerminus 2High11.5%15 / 1308.3%9 / 45

Total: 45 tasks, all rollouts

Completing a
research workflow.

MatFlowBench measures whether an LLM agent can complete a specified materials research task using the available tools and data, within the task’s resource limits.

Each task’s executable verifier compares results with numerical references, using tolerances set by the underlying physics. Other checks cover physical constraints from well-established scientific workflows. Additionally, some tasks require calculation files or measurement logs.

A pass means that the submitted result meets these checks. It does not by itself establish new materials discovery, complete physical validation, or sound reasoning at every step.

01 Interpret the objective02 Use tools and data03 Submit the result04 Check the result

Computational and
experimental workflows.

In silico25 tasks

Computational

Agents are given simulation software and a scientific goal. They plan and run the calculations themselves, much as a computational materials scientist would and must carry the work through to the property or conclusion the task asks for.

From measured data20 tasks

Experimental

Agents receive a set of target properties and access to a simulated lab, where every experiment returns real measured data. Each sample can be characterized on several instruments, and each instrument reveals a different set of properties. With a finite budget to spend, the agent decides which samples to prepare and which measurements are worth running, using what it learns to narrow the search until it finds a material that meets the requirements.

What the score means.

Each rollout passes or fails. The task’s executable verifier makes that call, and the leaderboard reports how often each model passes.

01

Numerical references

Results are compared with numerical references, using tolerances set by the underlying physics.

02

Physical constraints and evidence

Other checks cover physical constraints from well-established scientific workflows. Some tasks also require calculation files or measurement logs, and experimental tasks must reach a qualifying material within the measurement budget.

03

Pass rate and Pass@1

Each model ran with high reasoning effort, twice per computational task and four times per experimental task. Pass rate is the share of all rollouts that passed, and the leaderboard is ranked by it. Pass@1 is the share of passing rollouts on a task, averaged so that every task counts equally.

To request an evaluation, contact science@collinear.ai.