LitXBench - Materials: leaderboard
Metric: Material F1 (0-1), whether the right set of materials is extracted for extracting every experiment (materials, process chains, microstructure configurations and measurements) as Python objects from the OCR text of LitXAlloy's 19 alloy papers (1,426 measurements; figures excluded), extracted materials matched to the annotated ones by the Hungarian algorithm, mean of three runs; models called through the API with Pydantic AI; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 0.97 |
| 2 | GPT-5.2 (High) | 0.97 |
| 3 | Gemini 3.1 Pro (Preview) | 0.96 |
| 4 | GPT-5 Mini | 0.94 |
| 5 | Claude Haiku 4.5 | 0.94 |
| 6 | Claude Opus 4.6 | 0.91 |
Interactive version: theaggregate.ai/benchmark?slug=litxbench-materials · How It Works · Data refreshed daily, snapshot 2026-10-07.