TopoBench (Medium): leaderboard

Metric: Accuracy (proportion of puzzles solved, times 100; the mean over the six families) on the medium tier, TopoBench's six topological grid-puzzle families (Flow Free, Bridges, Loopy, Galaxies, Undead, Pattern), 50 generated puzzles per family, shown as ASCII grids with one fixed worked example, reasoning at the highest available level, up to 100k tokens, single attempt; a puzzle counts as solved only if the family's constraint verifier accepts the returned JSON grid; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini (High)44#176 (GPT-5 Mini)
2DeepSeek V3.2 (Thinking)37#198 (DeepSeek V3.2)
3Gemini 3 Flash (Preview)35#78
4Qwen 3 235B A22B 2507 (Thinking)12#253 (Qwen 3 235B A22B 2507)
5OLMo 3.1 32B (Thinking)1#622 (OLMo 3.1 32B)
6Llama 4 Maverick0#451
7GLM-4.7 Flash0#496
8Qwen 3 32B (Thinking)0#424 (Qwen 3 32B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=topobench-medium · How It Works · Data refreshed daily, snapshot 2026-10-11.