AcuityBench (Conversational): leaderboard

Metric: Exact acuity match (%) in the conversational format on the 527 clear consensus cases: the model answers naturally without being asked for a label, a GPT-4.1 judge infers the communicated acuity level, mode of five samples at temperature 1; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash78
2GPT-5.477.2
3DeepSeek V3.176.8
4Gemini 2.5 Pro76.7
5Claude Sonnet 4.674.4
6Claude Opus 4.771.9
7GPT-4.171.4
8Claude Haiku 4.5 (20251001)70
9GPT-5 Mini67.7
10Qwen 2.5 72B Instruct Turbo66.8
11Llama 3.3 70B Instruct61.9

Interactive version: theaggregate.ai/benchmark?slug=acuitybench-conversational · How It Works · Data refreshed daily, snapshot 2026-10-07.