MentalBench (DSM Diagnosis) - Resolvable Differential Diagnosis: leaderboard

Metric: Accuracy (%) on the 13,050 Type 4 cases (cases that include the discriminating evidence, so only one diagnosis of the pair is correct; choose one or more); cases synthesized from a psychiatrist-built DSM-5 knowledge graph by three LLMs and validated by a psychiatrist and a clinical psychologist; four options per case; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini77.72#176
2Claude Sonnet 4.574.84#138
3GPT-5.170.3#131
4Claude Haiku 4.565#271
5Gemini 2.5 Flash60.96#237
6Gemini 2.5 Pro59.82#145
7GPT-4o48.3#333
8Llama 3.1 70B Instruct34.97#548
9Qwen 2.5 14B Instruct31.83#634
10Qwen 2.5 32B Instruct30.89#491
11Qwen 3 235B A22B28.63#304
12Qwen 3 32B27.37#424
13Gemma 3 4B (IT)18.41#971
14Qwen 2.5 72B Instruct18.25#436
15Gemma 3 12B (IT)15.79#655

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mentalbench-dsm-diagnosis-resolvable-differential-diagnosis · How It Works · Data refreshed daily, snapshot 2026-10-11.