MHGraphBench - Two-hop Verification with Evidence: leaderboard

Metric: Accuracy (%) on the 1,200 two-hop verification items with sanitized PrimeKG feature snippets added, letter-only answers; the OpenAI API models answer at temperature 0 (at most 120 completion tokens) with strict answer-letter parsing, the open models are scored by forced-choice option-letter log-probabilities; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1GPT-4.171.42
2GPT-4o68.83
3Qwen 2.5 32B Instruct61.25
4GPT-5.2 Instant61.08
5GPT-5.1 Instant60
6GPT-5 Mini59.67
7Qwen 2.5 7B Instruct56.75
8Mistral 7B Instruct (v0.3)56.17
9Llama 3.1 8B Instruct50.25
10DeepSeek-R1-Distill-Qwen-7B49.75
11DeepSeek R1 Distill Qwen 32B49.58

Interactive version: theaggregate.ai/benchmark?slug=mhgraphbench-two-hop-verification-with-evidence · How It Works · Data refreshed daily, snapshot 2026-10-07.