QCalEval (In-Context) - Significance Analysis: leaderboard

Metric: Score (0-100) for the scientific significance analysis, three expert key points per item scored 0, 0.5 or 1 by GPT-5.4 and Gemini 3.1 Pro Preview and averaged, with one analysis demonstration per scenario type, on the QCalEval quantum-calibration plots with in-context demonstrations from the same experiment family (scenario types with a single sample left out), greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1Claude Opus 4.684.7
2Gemini 3.1 Pro (Preview)81.3
3GPT-5.481
4Gemma 4 31B (IT)80.6
5Gemini 3.1 Flash Lite78.5
6Claude Sonnet 4.677.8
7Claude Haiku 4.566.1
8GPT-5.4 Mini58.8
9InternVL3-38B56.2
10InternVL3-78B50.5
11Qwen 3.5 27B41.8
12Qwen 3.5 397B A17B37.4
13Qwen 3.5 122B A10B36.1
14Qwen 3.5 35B A3B33.4
15Qwen 3.5 9B32.8

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-in-context-significance-analysis · How It Works · Data refreshed daily, snapshot 2026-10-07.