QCalEval - Technical Description: leaderboard

Metric: Score (0-100) for a structured JSON description of plot type, axes and salient features (half exact-match fields, half key-point coverage judged by GPT-5.4 and Gemini 3.1 Pro Preview), on the 243 quantum-calibration plots of QCalEval (22 experiment families, superconducting qubits and neutral atoms), zero-shot with the family background text, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1GPT-5.490.9
2Claude Opus 4.690.8
3GPT-5.4 Mini90.3
4Claude Sonnet 4.689.7
5Gemini 3.1 Flash Lite89.2
6Gemini 3.1 Pro (Preview)88.5
7Qwen 3.5 397B A17B88.1
8Qwen 3.5 27B87
9Qwen 3.5 35B A3B86.8
10Qwen 3.5 122B A10B86.6
11Gemma 4 31B (IT)85.6
12Claude Haiku 4.583.4
13Qwen 3.5 9B81.5
14InternVL3-38B79.2
15InternVL3-78B76.3

Interactive version: theaggregate.ai/benchmark?slug=qcaleval-technical-description · How It Works · Data refreshed daily, snapshot 2026-10-07.