LitXBench - Measurements: leaderboard

Metric: Measurement F1 (0-1), extracted measurement values against the annotated ones (kind, value, unit, uncertainty, temperature and pressure) for extracting every experiment (materials, process chains, microstructure configurations and measurements) as Python objects from the OCR text of LitXAlloy's 19 alloy papers (1,426 measurements; figures excluded), extracted materials matched to the annotated ones by the Hungarian algorithm, mean of three runs; models called through the API with Pydantic AI; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)0.71
2GPT-5.2 (High)0.65
3Claude Opus 4.60.62
4Gemini 3 Flash (Preview)0.61
5GPT-5 Mini0.52
6Claude Haiku 4.50.51

Interactive version: theaggregate.ai/benchmark?slug=litxbench-measurements · How It Works · Data refreshed daily, snapshot 2026-10-07.