MentalBench (DSM Diagnosis) - Incomplete Patient Narratives: leaderboard

Metric: Accuracy (%) on the 3,450 Type 2 cases (colloquial patient self-reports with diagnostic information left incomplete; single answer); cases synthesized from a psychiatrist-built DSM-5 knowledge graph by three LLMs and validated by a psychiatrist and a clinical psychologist; four options per case; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.591.16#138
2GPT-5.190.43#131
3GPT-5 Mini88.46#176
4Gemini 2.5 Pro88.2#145
5Claude Haiku 4.587.39#271
6Gemini 2.5 Flash85.27#237
7Qwen 3 235B A22B84.17#304
8Gemma 3 27B (IT)84.09#509
9Qwen 3 32B83.13#424
10Qwen 2.5 72B Instruct83.1#436
11Qwen 2.5 32B Instruct83.04#491
12Qwen 2.5 14B Instruct82.26#634
13Llama 3.1 70B Instruct80.81#548
14GPT-4o79.97#333
15Gemma 3 12B (IT)78.32#655

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mentalbench-dsm-diagnosis-incomplete-patient-narratives · How It Works · Data refreshed daily, snapshot 2026-10-11.