QCalEval - Parameter Extraction: leaderboard

Metric: Score (0-100) for family-specific physical parameters extracted as JSON, per-field tolerance scoring, failed parses scored 0, on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)71.5
2Gemma 4 31B (IT)68.3
3Claude Opus 4.664.7
4GPT-5.464.3
5Gemini 3.1 Flash Lite63.8
6GPT-5.4 Mini62.6
7Qwen 3.5 397B A17B62.5
8Qwen 3.5 122B A10B61.2
9Claude Sonnet 4.660.4
10Qwen 3.5 27B58.7
11Qwen 3.5 35B A3B57.8
12Qwen 3.5 9B57.1
13InternVL3-78B52.9
14Claude Haiku 4.551
15InternVL3-38B49.2

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-parameter-extraction · How It Works · Data refreshed daily, snapshot 2026-10-07.