TopoBench (Hard): leaderboard
Metric: Accuracy (proportion of puzzles solved, times 100; the mean over the six families) on the hard tier, TopoBench's six topological grid-puzzle families (Flow Free, Bridges, Loopy, Galaxies, Undead, Pattern), 50 generated puzzles per family, shown as ASCII grids with one fixed worked example, reasoning at the highest available level, up to 100k tokens, single attempt; a puzzle counts as solved only if the family's constraint verifier accepts the returned JSON grid; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 Mini (High) | 24 | #176 (GPT-5 Mini) |
| 2 | DeepSeek V3.2 (Thinking) | 10 | #198 (DeepSeek V3.2) |
| 3 | Gemini 3 Flash (Preview) | 9 | #78 |
| 4 | Qwen 3 235B A22B 2507 (Thinking) | 1 | #253 (Qwen 3 235B A22B 2507) |
| 5 | GLM-4.7 Flash | 1 | #496 |
| 6 | OLMo 3.1 32B (Thinking) | 1 | #622 (OLMo 3.1 32B) |
| 7 | Llama 4 Maverick | 0 | #451 |
| 8 | Qwen 3 32B (Thinking) | 0 | #424 (Qwen 3 32B) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=topobench-hard · How It Works · Data refreshed daily, snapshot 2026-10-11.