AttuneBench - Pairwise Accuracy: leaderboard

Metric: Pairwise accuracy (%) at predicting the participant's preferences among three response variants per turn (original, preference-informed alternate, participant-written golden), Default mode (conversation only), 200 conversations; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.764.6
2Claude Opus 4.663.7
3MiMo-V2-Pro59.8
4Claude Haiku 4.556
5GPT-5.553.8
6Mistral Large51.7
7Claude Sonnet 4.651.2
8Gemini 3.1 Pro (Preview)50.3
9GPT-5.447
10Grok 446.1
11Qwen 2.5 72B45

Interactive version: theaggregate.ai/benchmark?slug=attunebench-pairwise-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.