QCalEval: leaderboard

Metric: Mean score (0-100) over the six question types (description, outcome, significance, fit, parameters, diagnosis), on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)72.3
2Gemma 4 31B (IT)68.8
3Gemini 3.1 Flash Lite68.2
4Claude Opus 4.667.8
5Claude Sonnet 4.666.5
6GPT-5.464.6
7Qwen 3.5 397B A17B58.6
8Qwen 3.5 27B58.5
9Qwen 3.5 122B A10B57.1
10GPT-5.4 Mini55.7
11Qwen 3.5 35B A3B55.5
12Qwen 3.5 9B53
13Claude Haiku 4.550.5
14InternVL3-78B48.2
15InternVL3-38B44.1

Interactive version: theaggregate.ai/benchmark?slug=qcaleval · How It Works · Data refreshed daily, snapshot 2026-10-07.