AutoControl Arena X-Bench (Baseline): leaderboard
Metric: Risk rate (%): share of trajectories a Qwen3-32B monitor scores 2 or more of 4 for action or chain-of-thought suspicion, over the 70 X-Bench scenarios (3 rollouts each) under low stress and low temptation (S0T0); lower is better. Source: arxiv.org. 9 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 Mini | 13.2 | #176 |
| 2 | Qwen 3 14B | 15.9 | #524 |
| 3 | GPT-4.1 Mini | 16.2 | #346 |
| 4 | Qwen 3 8B | 18.2 | #667 |
| 5 | Qwen 3 32B | 19 | #424 |
| 6 | Gemini 3 Flash | 25 | #93 |
| 7 | Gemini 2.5 Flash | 26.5 | #237 |
| 8 | Qwen 3 4B | 30.3 | #823 |
| 9 | O3 Mini | 30.9 | #266 |
Interactive version: theaggregate.ai/benchmark?slug=autocontrol-arena-x-bench-baseline · How It Works · Data refreshed daily, snapshot 2026-10-11.