MHGraphBench - Two-hop Verification: leaderboard

Metric: Accuracy (%) on the 1,200 two-hop verification items (about half yes), letter-only answers; the OpenAI API models answer at temperature 0 (at most 120 completion tokens) with strict answer-letter parsing, the open models are scored by forced-choice option-letter log-probabilities; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 15 models tracked.

Top models

#ModelScore
1GPT-4o62.08
2GPT-4.157.58
3Qwen 2.5 32B Instruct50.5
4Llama 3.1 8B Instruct50.42
5GPT-5 Mini50.25
6GPT-5.2 Instant50.17
7GPT-5.1 Instant50.17
8Qwen 2.5 7B Instruct50.08
9DeepSeek R1 Distill Qwen 32B49.75
10DeepSeek-R1-Distill-Qwen-7B49.75
11Mistral 7B Instruct (v0.3)49.33

Interactive version: theaggregate.ai/benchmark?slug=mhgraphbench-two-hop-verification · How It Works · Data refreshed daily, snapshot 2026-10-07.