Blind-Spots-Bench - Text-only (Code Execution): leaderboard
Metric: Accuracy (%; mean@4 on the text-only questions with a Python execution tool allowing up to five tool calls per question; 235 human-authored questions with reference solutions graded correct or incorrect by gemini-3-flash with code execution; thinking enabled at medium effort where available; max 32,768 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) (Medium) | 86.11 |
| 2 | GLM-5.1 | 75 |
| 3 | GLM-5.2 | 75 |
| 4 | GPT-5.4 (Medium) | 73.61 |
| 5 | Kimi K2.6 | 71.3 |
| 6 | Qwen 3.5 397B A17B | 68.98 |
| 7 | Kimi K2.5 | 68.06 |
| 8 | GPT-OSS-120B | 67.13 |
| 9 | GPT-5.4 Mini (Medium) | 66.67 |
| 10 | Qwen 3 VL 30B A3B (Thinking) | 56.94 |
Interactive version: theaggregate.ai/benchmark?slug=blind-spots-bench-text-only-code-execution · How It Works · Data refreshed daily, snapshot 2026-09-29.