Blind-Spots-Bench - Text-only (Code Execution): leaderboard

Metric: Accuracy (%; mean@4 on the text-only questions with a Python execution tool allowing up to five tool calls per question; 235 human-authored questions with reference solutions graded correct or incorrect by gemini-3-flash with code execution; thinking enabled at medium effort where available; max 32,768 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (Medium)86.11
2GLM-5.175
3GLM-5.275
4GPT-5.4 (Medium)73.61
5Kimi K2.671.3
6Qwen 3.5 397B A17B68.98
7Kimi K2.568.06
8GPT-OSS-120B67.13
9GPT-5.4 Mini (Medium)66.67
10Qwen 3 VL 30B A3B (Thinking)56.94

Interactive version: theaggregate.ai/benchmark?slug=blind-spots-bench-text-only-code-execution · How It Works · Data refreshed daily, snapshot 2026-09-29.