MHGraphBench - Two-hop Selection with Evidence: leaderboard

Metric: Accuracy (%) on the 1,200 two-hop selection items with sanitized PrimeKG feature snippets added, letter-only answers; the OpenAI API models answer at temperature 0 (at most 120 completion tokens) with strict answer-letter parsing, the open models are scored by forced-choice option-letter log-probabilities; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1GPT-5.2 Instant67.58
2GPT-5 Mini65.42
3GPT-5.1 Instant64.42
4GPT-4.161.5
5GPT-4o61.42
6Qwen 2.5 32B Instruct50.08
7Qwen 2.5 7B Instruct37.75
8DeepSeek R1 Distill Qwen 32B35.75
9Llama 3.1 8B Instruct31.08
10Mistral 7B Instruct (v0.3)28.25
11DeepSeek-R1-Distill-Qwen-7B25

Interactive version: theaggregate.ai/benchmark?slug=mhgraphbench-two-hop-selection-with-evidence · How It Works · Data refreshed daily, snapshot 2026-10-07.