LogicGraph - Precision: leaderboard

Metric: Precision (%): share of the model's generated proof paths that are valid (LogicGraph: 900 natural-language first-order logic problems generated from symbolic proof DAGs with 2 to 19 valid minimal derivation paths (300 each with 2-4, 5-7 and 8 or more paths), every model prompted to write as many independent proofs as possible; each step is auto-formalized by DeepSeek-V3.2-Exp and verified with Prover9; mean of the three tiers); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)90.1#64
2DeepSeek V3.2 Exp (Thinking)76.96#227 (DeepSeek V3.2 Exp)
3Qwen 3 235B A22B (Thinking)76.89#304 (Qwen 3 235B A22B)
4Gemini 2.5 Pro73.38#145
5Gemini 2.5 Flash67.04#237
6DeepSeek V3.2 Exp (Non-reasoning)52.04#227 (DeepSeek V3.2 Exp)
7O4 Mini51.38#172
8O350.79#121
9Kimi K2 (Thinking)48.2#236 (Kimi K2)
10Claude Sonnet 4.5 (Thinking)47.82#138 (Claude Sonnet 4.5)
11Claude Sonnet 4.536.05#138
12QwQ-32B25.91#410
13Qwen 3 235B A22B (Non-reasoning)13.89#304 (Qwen 3 235B A22B)
14GPT-OSS-120B9.97#330
15GPT-5.19.69#131

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicgraph-precision · How It Works · Data refreshed daily, snapshot 2026-10-11.