TopoBench (Hard): leaderboard

Metric: Accuracy (proportion of puzzles solved, times 100; the mean over the six families) on the hard tier, TopoBench's six topological grid-puzzle families (Flow Free, Bridges, Loopy, Galaxies, Undead, Pattern), 50 generated puzzles per family, shown as ASCII grids with one fixed worked example, reasoning at the highest available level, up to 100k tokens, single attempt; a puzzle counts as solved only if the family's constraint verifier accepts the returned JSON grid; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini (High)24#176 (GPT-5 Mini)
2DeepSeek V3.2 (Thinking)10#198 (DeepSeek V3.2)
3Gemini 3 Flash (Preview)9#78
4Qwen 3 235B A22B 2507 (Thinking)1#253 (Qwen 3 235B A22B 2507)
5GLM-4.7 Flash1#496
6OLMo 3.1 32B (Thinking)1#622 (OLMo 3.1 32B)
7Llama 4 Maverick0#451
8Qwen 3 32B (Thinking)0#424 (Qwen 3 32B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=topobench-hard · How It Works · Data refreshed daily, snapshot 2026-10-11.