QCalEval - Fit Assessment: leaderboard

Metric: Accuracy (%) of the three-way fit reliability label (reliable, unreliable, no fit), on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)84.4
2Gemma 4 31B (IT)82.7
3Gemini 3.1 Flash Lite82.7
4Claude Sonnet 4.676.5
5Claude Opus 4.676.1
6Qwen 3.5 27B56.4
7GPT-5.454.7
8Qwen 3.5 35B A3B52.7
9Qwen 3.5 397B A17B50.6
10Qwen 3.5 122B A10B50.2
11Qwen 3.5 9B49.8
12Claude Haiku 4.548.6
13InternVL3-78B42.8
14GPT-5.4 Mini42
15InternVL3-38B33.7

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-fit-assessment · How It Works · Data refreshed daily, snapshot 2026-10-07.