LitXBench - Composition Extraction (Code Output): leaderboard

Metric: F1 (0-1) for extracting the experimentally synthesized compositions from the OCR text of LitXAlloy's 19 alloy papers, mean of three runs; the model returns a pymatgen Composition object, with composition-normalizing helper functions provided; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)0.99
2Gemini 3.1 Pro (Preview)0.98
3GPT-5 Mini0.98
4Claude Opus 4.60.97
5GPT-5.2 (High)0.97
6Claude Haiku 4.50.77

Interactive version: theaggregate.ai/benchmark?slug=litxbench-composition-extraction-code-output · How It Works · Data refreshed daily, snapshot 2026-10-07.