MHGraphBench - Entity Clustering: leaderboard

Metric: Accuracy (%) on the 2,000 five-entity odd-one-out entity clustering items, letter-only answers; the OpenAI API models answer at temperature 0 (at most 120 completion tokens) with strict answer-letter parsing, the open models are scored by forced-choice option-letter log-probabilities; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-5.1 Instant92.35
2GPT-4o91.85
3GPT-4.191.85
4GPT-5 Mini91.75
5GPT-5.2 Instant90.1
6Qwen 2.5 32B Instruct54.6
7Qwen 2.5 7B Instruct36.95
8Mistral 7B Instruct (v0.3)28.15
9DeepSeek-R1-Distill-Qwen-7B20.4
10DeepSeek R1 Distill Qwen 32B18.9
11Llama 3.1 8B Instruct17.5

Interactive version: theaggregate.ai/benchmark?slug=mhgraphbench-entity-clustering · How It Works · Data refreshed daily, snapshot 2026-10-07.