MHGraphBench - Relation Prediction: leaderboard

Metric: Accuracy (%) on the 1,634 four-way drug-disease relation prediction items (indication, contraindication, off-label use, none), letter-only answers; the OpenAI API models answer at temperature 0 (at most 120 completion tokens) with strict answer-letter parsing, the open models are scored by forced-choice option-letter log-probabilities; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1GPT-5.2 Instant58.63
2GPT-5.1 Instant58.08
3GPT-5 Mini57.28
4GPT-4.154.96
5GPT-4o53.55
6Qwen 2.5 32B Instruct38.43
7DeepSeek-R1-Distill-Qwen-7B32.62
8Qwen 2.5 7B Instruct25.89
9Mistral 7B Instruct (v0.3)25.7
10Llama 3.1 8B Instruct23.5
11DeepSeek R1 Distill Qwen 32B22.46

Interactive version: theaggregate.ai/benchmark?slug=mhgraphbench-relation-prediction · How It Works · Data refreshed daily, snapshot 2026-10-07.