ATAD - Bridge Sentence Evaluation: leaderboard

Metric: Accuracy (%) on the T4 bridge sentence evaluation task (pick the incoherent one of five bridge sentences) of ATAD's text anomaly detection problems, mean over four datasets of 100 problems per task generated by GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet and LLaMA-3.3-70B teacher-student-orchestrator agents; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.3 70B Instruct60#520
2Claude 3.5 Sonnet59.5#337
3Gemini 2.0 Flash58.25#331
4GPT-4o Mini54#588
5GPT-4o53.25#333
6O4 Mini53#172
7Gemini 2.0 Flash Lite52.25#438
8Claude 3 Haiku51.75#784
9Gemini 1.5 Flash48.75#619
10GPT-3.5 Turbo48.5#849
11Claude 3.5 Haiku5#553

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=atad-bridge-sentence-evaluation · How It Works · Data refreshed daily, snapshot 2026-10-11.