MindBench - Alignment: leaderboard

Metric: Alignment with human consensus (0-1, 1 - MAE / 3, higher is better). Source: mindbench.ai. 18 models tracked.

Top models

#ModelScore
1O30.83
2GPT-5.50.75
3Grok 4.30.74
4Claude Haiku 4.5 (20251001)0.73
5GPT-5.40.72
6Mistral Medium 3.50.72
7DeepSeek V4 Flash0.7
8Sonar0.7
9DeepSeek V4 Pro0.69
10Claude Fable 50.68
11Claude Opus 4.70.67
12Claude Sonnet 4.60.66
13Gemini 3.5 Flash0.66
14Mistral Small 40.66
15Grok 4.1 Fast0.65

Interactive version: theaggregate.ai/benchmark?slug=mindbench-alignment · How It Works · Data refreshed daily, snapshot 2026-09-19.