ATAD - Tone and Style Violation: leaderboard
Metric: Accuracy (%) on the T7 tone and style violation task (pick the sentence that breaks tone or register) of ATAD's text anomaly detection problems, mean over four datasets of 100 problems per task generated by GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet and LLaMA-3.3-70B teacher-student-orchestrator agents; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.0 Flash | 88 | #331 |
| 2 | Claude 3.5 Sonnet | 86.75 | #337 |
| 3 | Gemini 2.0 Flash Lite | 86.25 | #438 |
| 4 | Llama 3.3 70B Instruct | 84.25 | #520 |
| 5 | GPT-4o Mini | 83 | #588 |
| 6 | GPT-3.5 Turbo | 81.5 | #849 |
| 7 | GPT-4o | 81 | #333 |
| 8 | O4 Mini | 80 | #172 |
| 9 | Claude 3 Haiku | 72.75 | #784 |
| 10 | Claude 3.5 Haiku | 35.5 | #553 |
| 11 | Gemini 1.5 Flash | 21 | #619 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=atad-tone-and-style-violation · How It Works · Data refreshed daily, snapshot 2026-10-11.