MHGraphBench - Entity Typing: leaderboard

Metric: Accuracy (%) on the 1,847 five-way entity typing items, letter-only answers; the OpenAI API models answer at temperature 0 (at most 120 completion tokens) with strict answer-letter parsing, the open models are scored by forced-choice option-letter log-probabilities; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-5 Mini98.48
2GPT-5.1 Instant98.27
3GPT-5.2 Instant98.05
4GPT-4.197.62
5GPT-4o97.4
6Qwen 2.5 32B Instruct76.45
7Mistral 7B Instruct (v0.3)70.76
8Qwen 2.5 7B Instruct49.32
9DeepSeek-R1-Distill-Qwen-7B9.42
10Llama 3.1 8B Instruct8.39
11DeepSeek R1 Distill Qwen 32B6.55

Interactive version: theaggregate.ai/benchmark?slug=mhgraphbench-entity-typing · How It Works · Data refreshed daily, snapshot 2026-10-07.