OpenAI MentalHealthBench: leaderboard

OpenAI benchmark of 1,215 synthetic mental health conversations, from everyday stress to emergencies, with adult, teen, caregiver and clinician users in 11 languages. GPT-5.6 Sol grades the reply to the last message against rubrics written by licensed psychologists and psychiatrists; score is the mean task-clipped share of positive rubric points.

Metric: Task-clipped rubric score (%). Source: openai.com. Saturation forecast: Around 2031. 17 models tracked.

Top models

#ModelScore
1GPT-657.33
2GPT-6 Sol53.94
3Claude Opus 5.552.38
4GPT-6 Luna50.21
5Muse Spark 1.348.6
6GPT-5.6 Sol (August 2026)46.98
7Claude Fable 5.146.36
8GPT-5.6 Luna (August 2026)44.89
9Claude Sonnet 544.54
10GPT-5 (Thinking)42.9
11Claude Haiku 4.541.73
12Grok 4.741.3
13Gemini 3.8 Flash35.5
14Gemini 2.5 Flash33.5
15GPT-4o (Mar 2025)32.08

Interactive version: theaggregate.ai/benchmark?slug=openai-mentalhealthbench · How It Works · Data refreshed daily, snapshot 2026-09-24.