OpenAI MentalHealthBench - Behavior - Context Seeking: leaderboard

Metric: Axis score, normalized to its possible range (%). Source: openai.com. Saturation forecast: Around 2029. 15 models tracked.

Top models

#ModelScore
1Claude Opus 5.556.88
2GPT-652.83
3Muse Spark 1.352.26
4Claude Fable 5.148.62
5GPT-6 Sol47.43
6Claude Sonnet 546.11
7GPT-5.6 Sol (August 2026)36.48
8GPT-6 Luna36.24
9Claude Haiku 4.535.91
10GPT-5.6 Luna (August 2026)34.16
11Gemini 3.8 Flash25.39
12Grok 4.724.31
13Gemini 3.1 Pro (Preview)22.11
14GPT-4o (Mar 2025)18.72
15Gemini 2.5 Pro15.32

Interactive version: theaggregate.ai/benchmark?slug=openai-mentalhealthbench-behavior-context-seeking · How It Works · Data refreshed daily, snapshot 2026-09-24.