QCalEval - Outcome Classification: leaderboard

Metric: Accuracy (%) of the four-way experimental outcome label (expected behavior, suboptimal parameters, anomalous behavior, apparatus issue), on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 18 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)57.2
2Gemma 4 31B (IT)54.3
3Gemini 3.1 Flash Lite53.5
4GPT-5.452.7
5Claude Opus 4.649
6Claude Sonnet 4.648.6
7Qwen 3.5 27B45.7
8Qwen 3.5 122B A10B44
9Qwen 3.5 397B A17B42.8
10Qwen 3.5 35B A3B39.9
11GPT-5.4 Mini39.5
12Qwen 3.5 9B37.9
13InternVL3-78B37
14Claude Haiku 4.536.6
15InternVL3-38B34.6

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-outcome-classification · How It Works · Data refreshed daily, snapshot 2026-10-07.