AttuneBench (Verbose): leaderboard

Metric: Composite (0-100) = 100 x (0.24 emotion tracking + 0.49 turn-level verification + 0.27 holistic comprehension), the evaluated model stepping through real human-model conversations turn by turn against the participants' own annotations, Verbose mode (the model also writes out its reasoning), 50-conversation subsample; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.753.95
2Claude Opus 4.653.5
3GPT-5.553.44
4MiMo-V2-Pro52.71
5Qwen 2.5 72B51.48
6Claude Haiku 4.551.47
7Gemini 3.1 Pro (Preview)51.21
8Claude Sonnet 4.649.73
9Grok 449.7
10Mistral Large49.66
11GPT-5.449.1

Interactive version: theaggregate.ai/benchmark?slug=attunebench-verbose · How It Works · Data refreshed daily, snapshot 2026-10-07.