QCalEval - Calibration Diagnosis: leaderboard

Metric: Accuracy (%) of the family-specific calibration status code, on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)71.2
2Gemma 4 31B (IT)62.1
3GPT-5.461.3
4Gemini 3.1 Flash Lite60.9
5Claude Opus 4.660.5
6Claude Sonnet 4.660.1
7Qwen 3.5 397B A17B55.6
8Qwen 3.5 27B55.1
9Qwen 3.5 9B52.3
10Qwen 3.5 122B A10B51.9
11GPT-5.4 Mini51.4
12Qwen 3.5 35B A3B50.6
13InternVL3-78B45.7
14Claude Haiku 4.542.8
15InternVL3-38B40.3

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-calibration-diagnosis · How It Works · Data refreshed daily, snapshot 2026-10-07.