ATAD - Bridge Sentence Evaluation: leaderboard
Metric: Accuracy (%) on the T4 bridge sentence evaluation task (pick the incoherent one of five bridge sentences) of ATAD's text anomaly detection problems, mean over four datasets of 100 problems per task generated by GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet and LLaMA-3.3-70B teacher-student-orchestrator agents; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Llama 3.3 70B Instruct | 60 | #520 |
| 2 | Claude 3.5 Sonnet | 59.5 | #337 |
| 3 | Gemini 2.0 Flash | 58.25 | #331 |
| 4 | GPT-4o Mini | 54 | #588 |
| 5 | GPT-4o | 53.25 | #333 |
| 6 | O4 Mini | 53 | #172 |
| 7 | Gemini 2.0 Flash Lite | 52.25 | #438 |
| 8 | Claude 3 Haiku | 51.75 | #784 |
| 9 | Gemini 1.5 Flash | 48.75 | #619 |
| 10 | GPT-3.5 Turbo | 48.5 | #849 |
| 11 | Claude 3.5 Haiku | 5 | #553 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=atad-bridge-sentence-evaluation · How It Works · Data refreshed daily, snapshot 2026-10-11.