QEncodeBench - 3-Coloring: leaderboard

Metric: Semantic pass@1 (%; L3 gate, graph 3-coloring (all edges bichromatic, surjective 2-bit color code), 100 instances of the core set; the model writes a Qiskit build_oracle function in one shot and the verifier decides full solution-set equivalence up to a global phase by exhaustive simulation, with ancillas restored and a measurement-free unitary circuit; pass@1, one attempt per instance (five samples at temperature 0.7 for the non-reasoning DeepSeek-V4-Flash row)). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-OSS-120B (High)60
2Claude Opus 4.8 (Claude Code)49
3GPT-OSS-20B (High)43
4DeepSeek V4 Flash (Thinking)25
5Claude Haiku 4.5 (Claude Code)17
6DeepSeek V4 Flash (Non-reasoning)2.4
7DeepSeek R1 0528 Qwen3 8B0
8Qwen 2.5 Coder 7B Instruct0

Interactive version: theaggregate.ai/benchmark?slug=qencodebench-3-coloring · How It Works · Data refreshed daily, snapshot 2026-09-26.