QCalEval - Significance Analysis: leaderboard

Metric: Score (0-100) for the scientific significance analysis, three expert key points per item scored 0, 0.5 or 1 by GPT-5.4 and Gemini 3.1 Pro Preview and averaged, on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.665.5
2GPT-5.463.7
3Claude Sonnet 4.663.4
4Gemini 3.1 Pro (Preview)61.1
5Gemma 4 31B (IT)59.8
6Gemini 3.1 Flash Lite59.4
7Qwen 3.5 397B A17B52
8Qwen 3.5 122B A10B49
9Qwen 3.5 27B48.3
10GPT-5.4 Mini48.3
11Qwen 3.5 35B A3B45.7
12Claude Haiku 4.540.8
13Qwen 3.5 9B39.5
14InternVL3-78B34.1
15InternVL3-38B27.6

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-significance-analysis · How It Works · Data refreshed daily, snapshot 2026-10-07.