QEncodeBench - String Matching: leaderboard
Metric: Semantic pass@1 (%; L3 gate, string matching (the pattern occurs at some offset), 70 instances of the core set; the model writes a Qiskit build_oracle function in one shot and the verifier decides full solution-set equivalence up to a global phase by exhaustive simulation, with ancillas restored and a measurement-free unitary circuit; pass@1, one attempt per instance (five samples at temperature 0.7 for the non-reasoning DeepSeek-V4-Flash row)). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-OSS-120B (High) | 90 |
| 2 | Claude Opus 4.8 (Claude Code) | 87.1 |
| 3 | GPT-OSS-20B (High) | 75.7 |
| 4 | Claude Haiku 4.5 (Claude Code) | 68.6 |
| 5 | DeepSeek V4 Flash (Thinking) | 61.4 |
| 6 | DeepSeek V4 Flash (Non-reasoning) | 24.9 |
| 7 | DeepSeek R1 0528 Qwen3 8B | 4.3 |
| 8 | Qwen 2.5 Coder 7B Instruct | 0 |
Interactive version: theaggregate.ai/benchmark?slug=qencodebench-string-matching · How It Works · Data refreshed daily, snapshot 2026-09-26.