LitXBench (Coding Agents): leaderboard

Metric: Overall F1 (0-1), the weighted combination of measurement (0.5), process condition (0.2), material set (0.15) and configuration (0.15) F1 for extracting every experiment (materials, process chains, microstructure configurations and measurements) as Python objects from the OCR text of LitXAlloy's 19 alloy papers (1,426 measurements; figures excluded), extracted materials matched to the annotated ones by the Hungarian algorithm, mean of three runs; coding agents (Claude Code, Codex, Gemini CLI); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)0.8
2Claude Opus 4.60.78
3GPT-5.2 Codex (High)0.73

Interactive version: theaggregate.ai/benchmark?slug=litxbench-coding-agents · How It Works · Data refreshed daily, snapshot 2026-10-07.