QuanBench+ (Pass@5) - PennyLane: leaderboard

Metric: Pass@5 (%): any of five PennyLane programs sampled at temperature 0.8 passes, on the 42 framework-aligned QuanBench+ tasks (quantum algorithms, gate decomposition, state preparation), graded by executable functional tests with KL-divergence acceptance for probabilistic outputs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1DeepSeek R159.5
2GPT-5.157.1
3MiniMax-M2.154.8
4Kimi K2 (Thinking)52.4
5GLM-4.747.6
6Gemini 3 Pro40.5
7Claude 3.7 Sonnet38.1
8Gemini 2.5 Flash35.7
9GPT-4.135.7
10Llama 4 Maverick31
11Qwen 2.5 7B Instruct21.4

Interactive version: theaggregate.ai/benchmark?slug=quanbench-plus-pass-5-pennylane · How It Works · Data refreshed daily, snapshot 2026-10-07.