ATBench - Risk Source Diagnosis: leaderboard

Metric: Fine-grained diagnosis accuracy (%) of the risk-source label (ATBench taxonomy) over the 497 unsafe ATBench trajectories, with the AgentDoG diagnosis template; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 13 models tracked.

Top models

#ModelScore
1GPT-5.433.6
2GPT-5.229.5
3Gemini 3.1 Pro (Preview)24.8
4Gemini 3 Flash18.4
5QwQ-32B15.8
6Qwen 3.5 397B A17B7.7
7Qwen 3 235B A22B 2507 Instruct7
8Qwen 3.5 4B6.6
9Llama 3.1 8B Instruct6.2
10Qwen 2.5 7B Instruct5.3
11Qwen 3 4B4.4
12Qwen 3 4B 2507 Instruct1

Interactive version: theaggregate.ai/benchmark?slug=atbench-risk-source-diagnosis · How It Works · Data refreshed daily, snapshot 2026-10-07.