QuanBench+ (Feedback Repair) - Qiskit: leaderboard

Metric: Pass@1 after feedback repair (%): after a runtime error the model receives the exception trace, after a wrong answer its failing function, and returns a corrected Qiskit program, up to five repair attempts per task, on the 42 framework-aligned QuanBench+ tasks (quantum algorithms, gate decomposition, state preparation), graded by executable functional tests with KL-divergence acceptance for probabilistic outputs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1GPT-5.183.3
2Gemini 3 Pro73.8
3DeepSeek R171.4
4GLM-4.769
5Kimi K2 (Thinking)69
6Gemini 2.5 Flash61.9
7GPT-4.157.1
8Claude 3.7 Sonnet57.1
9MiniMax-M2.157.1
10Llama 4 Maverick54.8
11Qwen 2.5 7B Instruct19

Interactive version: theaggregate.ai/benchmark?slug=quanbench-plus-feedback-repair-qiskit · How It Works · Data refreshed daily, snapshot 2026-10-07.