MentalBench (DSM Diagnosis) - Complete Clinical Summaries: leaderboard

Metric: Accuracy (%) on the 1,725 Type 1 cases (structured professional summaries that meet all DSM-5 criteria; single answer); cases synthesized from a psychiatrist-built DSM-5 knowledge graph by three LLMs and validated by a psychiatrist and a clinical psychologist; four options per case; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.194.38#131
2Claude Sonnet 4.594.32#138
3Claude Haiku 4.594.2#271
4Qwen 2.5 14B Instruct94.15#634
5Gemini 2.5 Pro93.91#145
6Qwen 2.5 72B Instruct93.85#436
7Gemini 2.5 Flash93.39#237
8Gemma 3 27B (IT)92.87#509
9Qwen 2.5 32B Instruct92.87#491
10Gemma 3 12B (IT)92.81#655
11Llama 3.1 70B Instruct91.94#548
12Qwen 3 14B90.78#524
13Qwen 3 32B90.2#424
14Qwen 3 8B89.74#667
15Qwen 2.5 7B Instruct85.85#846

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mentalbench-dsm-diagnosis-complete-clinical-summaries · How It Works · Data refreshed daily, snapshot 2026-10-11.