QuanBench+ - PennyLane: leaderboard

Metric: Pass@1 (%): greedy decoding at temperature 0, one PennyLane program per task, on the 42 framework-aligned QuanBench+ tasks (quantum algorithms, gate decomposition, state preparation), graded by executable functional tests with KL-divergence acceptance for probabilistic outputs; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 12 models tracked.

Top models

#ModelScore
1GPT-5.142.9
2Gemini 3 Pro40.5
3DeepSeek R133.3
4MiniMax-M2.131
5Claude 3.7 Sonnet26.2
6Kimi K2 (Thinking)26.2
7Gemini 2.5 Flash23.8
8GPT-4.123.8
9GLM-4.723.8
10Llama 4 Maverick19
11Qwen 2.5 7B Instruct11.9

Interactive version: theaggregate.ai/benchmark?slug=quanbench-plus-pennylane · How It Works · Data refreshed daily, snapshot 2026-10-07.