Blind-Spots-Bench - Text-only: leaderboard

Metric: Accuracy (%; mean@4 over four samples on the text-only questions; 235 human-authored questions with reference solutions graded correct or incorrect by gemini-3-flash with code execution; no tools; thinking enabled at medium effort where available; max 32,768 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 31 models tracked.

Top models

#ModelScore
1GPT-5.584
2Gemini 3.1 Pro (Preview) (Medium)83.3
3GPT-5.4 (Medium)78.9
4Gemini 3 Flash (Preview) (Medium)76.2
5GLM-5.273.8
6Qwen 3.5 397B A17B71.1
7GPT-570.6
8GPT-5.2 (Medium)70.6
9DeepSeek V4 Pro (Reasoning)70.4
10Kimi K2.568.3
11GPT-5 Mini67.6
12DeepSeek V4 Flash (Reasoning)67.6
13GLM-5.167.4
14Kimi K2.665
15GLM-564.4

Interactive version: theaggregate.ai/benchmark?slug=blind-spots-bench-text-only · How It Works · Data refreshed daily, snapshot 2026-09-29.