RuleWorld - Multi-Hop (First-Order Logic): leaderboard

Metric: Exact-match accuracy (%; x100 of the 0-1 score; multi-hop rule QA with 2 to 4 hops, mean of the three hop counts, rules in first-order logic; all 1,000 abstract, non-commonsense procedural rules of the pool prepended to the prompt (full-context prompting); exact match on the boxed answer list, partial credit per matching position; 10 questions per task type and difficulty level for each of five seeds; temperature 0 (Qwen3-32B 0.6)). Source: arxiv.org. Saturation forecast: Around July 2028. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.651.33
2GPT-5.535.33
3DeepSeek V3.2 (Non-reasoning)34.67
4GPT-5.132
5Qwen 3 32B26
6Qwen 2.5 72B Instruct25.33
7Llama 3.1 70B Instruct20.67
8Qwen 2.5 32B Instruct20.67
9Qwen 2.5 14B Instruct19.33
10Llama 3.3 70B Instruct16.67

Interactive version: theaggregate.ai/benchmark?slug=ruleworld-multi-hop-first-order-logic · How It Works · Data refreshed daily, snapshot 2026-09-26.