LogicGraph - Family Recall: leaderboard

Metric: Family recall (%): share of the ground-truth families of derivation paths (paths sharing their inference nodes) the model reaches (the paper's Versatility) (LogicGraph: 900 natural-language first-order logic problems generated from symbolic proof DAGs with 2 to 19 valid minimal derivation paths (300 each with 2-4, 5-7 and 8 or more paths), every model prompted to write as many independent proofs as possible; each step is auto-formalized by DeepSeek-V3.2-Exp and verified with Prover9; mean of the three tiers); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)79.44#64
2Gemini 2.5 Pro69.13#145
3Qwen 3 235B A22B (Thinking)69.12#304 (Qwen 3 235B A22B)
4Gemini 2.5 Flash61.43#237
5DeepSeek V3.2 Exp (Thinking)61.02#227 (DeepSeek V3.2 Exp)
6Kimi K2 (Thinking)44.94#236 (Kimi K2)
7DeepSeek V3.2 Exp (Non-reasoning)31.38#227 (DeepSeek V3.2 Exp)
8Claude Sonnet 4.5 (Thinking)30.15#138 (Claude Sonnet 4.5)
9O330.02#121
10O4 Mini29.2#172
11QwQ-32B27.38#410
12Claude Sonnet 4.525.42#138
13Qwen 3 235B A22B (Non-reasoning)14.65#304 (Qwen 3 235B A22B)
14GPT-5.19.96#131
15GLM-4.69.66#246

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=logicgraph-family-recall · How It Works · Data refreshed daily, snapshot 2026-10-11.