Quantum API Drift: leaderboard

Metric: Mean version fidelity (%): Pass@1 on the prompted Qiskit version, averaged over the Qiskit 0.43, 1.3 and 2.0 targets (unbiased estimator from 3 samples per task over 50 Qiskit HumanEval-derived tasks); REST API access at temperature 0.8 with a 1,024-token output cap; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.

Top models

#ModelScore
1Claude Opus 4.766.89
2GPT-5.3 Codex60.1
3Grok 4.2057.78
4GPT-5.455.33
5Claude Sonnet 4.650.67
6Kimi K249.56
7Nemotron 3 Super44.52
8GPT-OSS-120B43
9Gemma 4 31B (IT)34.44
10Qwen3 Coder33.78
11Devstral 232.93
12GPT-5.4 Nano32.89
13Qwen 3 32B28.44
14GLM-4.7 Flash22.89
15Llama 3.3 70B Instruct17.33

Interactive version: theaggregate.ai/benchmark?slug=quantum-api-drift · How It Works · Data refreshed daily, snapshot 2026-09-29.