TopoBench (Easy): leaderboard

Metric: Accuracy (proportion of puzzles solved, times 100; the mean over the six families) on the easy tier, TopoBench's six topological grid-puzzle families (Flow Free, Bridges, Loopy, Galaxies, Undead, Pattern), 50 generated puzzles per family, shown as ASCII grids with one fixed worked example, reasoning at the highest available level, up to 100k tokens, single attempt; a puzzle counts as solved only if the family's constraint verifier accepts the returned JSON grid; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini (High)71#176 (GPT-5 Mini)
2Gemini 3 Flash (Preview)60#78
3DeepSeek V3.2 (Thinking)58#198 (DeepSeek V3.2)
4Qwen 3 235B A22B 2507 (Thinking)31#253 (Qwen 3 235B A22B 2507)
5Qwen 3 32B (Thinking)7#424 (Qwen 3 32B)
6OLMo 3.1 32B (Thinking)7#622 (OLMo 3.1 32B)
7GLM-4.7 Flash3#496
8Llama 4 Maverick0#451

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=topobench-easy · How It Works · Data refreshed daily, snapshot 2026-10-11.