QuanBench+ (Feedback Repair) - Cirq: leaderboard

Metric: Pass@1 after feedback repair (%): after a runtime error the model receives the exception trace, after a wrong answer its failing function, and returns a corrected Cirq program, up to five repair attempts per task, on the 42 framework-aligned QuanBench+ tasks (quantum algorithms, gate decomposition, state preparation), graded by executable functional tests with KL-divergence acceptance for probabilistic outputs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1Gemini 3 Pro76.2
2GPT-5.173.8
3DeepSeek R161.9
4GLM-4.761.9
5Claude 3.7 Sonnet59.5
6GPT-4.157.1
7Kimi K2 (Thinking)57.1
8Gemini 2.5 Flash50
9MiniMax-M2.147.6
10Llama 4 Maverick42.9
11Qwen 2.5 7B Instruct7.1

Interactive version: theaggregate.ai/benchmark?slug=quanbench-plus-feedback-repair-cirq · How It Works · Data refreshed daily, snapshot 2026-10-07.