QuanBench+ (Feedback Repair) - PennyLane: leaderboard

Metric: Pass@1 after feedback repair (%): after a runtime error the model receives the exception trace, after a wrong answer its failing function, and returns a corrected PennyLane program, up to five repair attempts per task, on the 42 framework-aligned QuanBench+ tasks (quantum algorithms, gate decomposition, state preparation), graded by executable functional tests with KL-divergence acceptance for probabilistic outputs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1GPT-5.166.7
2Gemini 3 Pro66.7
3GLM-4.764.3
4DeepSeek R152.4
5Claude 3.7 Sonnet47.6
6MiniMax-M2.147.6
7GPT-4.145.2
8Gemini 2.5 Flash40.5
9Llama 4 Maverick38.1
10Kimi K2 (Thinking)38.1
11Qwen 2.5 7B Instruct19

Interactive version: theaggregate.ai/benchmark?slug=quanbench-plus-feedback-repair-pennylane · How It Works · Data refreshed daily, snapshot 2026-10-07.