ATAD - Tone and Style Violation: leaderboard

Metric: Accuracy (%) on the T7 tone and style violation task (pick the sentence that breaks tone or register) of ATAD's text anomaly detection problems, mean over four datasets of 100 problems per task generated by GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet and LLaMA-3.3-70B teacher-student-orchestrator agents; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.0 Flash88#331
2Claude 3.5 Sonnet86.75#337
3Gemini 2.0 Flash Lite86.25#438
4Llama 3.3 70B Instruct84.25#520
5GPT-4o Mini83#588
6GPT-3.5 Turbo81.5#849
7GPT-4o81#333
8O4 Mini80#172
9Claude 3 Haiku72.75#784
10Claude 3.5 Haiku35.5#553
11Gemini 1.5 Flash21#619

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=atad-tone-and-style-violation · How It Works · Data refreshed daily, snapshot 2026-10-11.