LogicGraph: leaderboard

Metric: Success rate (%): share of problems with at least one fully valid proof path (LogicGraph: 900 natural-language first-order logic problems generated from symbolic proof DAGs with 2 to 19 valid minimal derivation paths (300 each with 2-4, 5-7 and 8 or more paths), every model prompted to write as many independent proofs as possible; each step is auto-formalized by DeepSeek-V3.2-Exp and verified with Prover9; mean of the three tiers); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)96.11#64
2Qwen 3 235B A22B (Thinking)91.22#304 (Qwen 3 235B A22B)
3DeepSeek V3.2 Exp (Thinking)88#227 (DeepSeek V3.2 Exp)
4Gemini 2.5 Pro87.78#145
5Gemini 2.5 Flash85.56#237
6Kimi K2 (Thinking)72.22#236 (Kimi K2)
7Claude Sonnet 4.5 (Thinking)61.67#138 (Claude Sonnet 4.5)
8DeepSeek V3.2 Exp (Non-reasoning)61#227 (DeepSeek V3.2 Exp)
9O358.33#121
10O4 Mini57.78#172
11Claude Sonnet 4.553.11#138
12QwQ-32B50.89#410
13Qwen 3 235B A22B (Non-reasoning)29.56#304 (Qwen 3 235B A22B)
14GPT-5.121.56#131
15GLM-4.621.11#246

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicgraph · How It Works · Data refreshed daily, snapshot 2026-10-11.