LogicGraph - Shortest Path Rate: leaderboard

Metric: Shortest-path finding rate (%): share of solutions matching the minimum ground-truth number of steps (LogicGraph: 900 natural-language first-order logic problems generated from symbolic proof DAGs with 2 to 19 valid minimal derivation paths (300 each with 2-4, 5-7 and 8 or more paths), every model prompted to write as many independent proofs as possible; each step is auto-formalized by DeepSeek-V3.2-Exp and verified with Prover9; mean of the three tiers); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)80.33#64
2Gemini 2.5 Pro72.78#145
3Qwen 3 235B A22B (Thinking)70.22#304 (Qwen 3 235B A22B)
4DeepSeek V3.2 Exp (Thinking)66.33#227 (DeepSeek V3.2 Exp)
5Gemini 2.5 Flash65.89#237
6Kimi K2 (Thinking)51.44#236 (Kimi K2)
7O339.44#121
8QwQ-32B35.44#410
9DeepSeek V3.2 Exp (Non-reasoning)34.44#227 (DeepSeek V3.2 Exp)
10Claude Sonnet 4.5 (Thinking)34.22#138 (Claude Sonnet 4.5)
11O4 Mini33.67#172
12Claude Sonnet 4.529.22#138
13Qwen 3 235B A22B (Non-reasoning)17.33#304 (Qwen 3 235B A22B)
14GLM-4.614#246
15GPT-5.113.56#131

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicgraph-shortest-path-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.