HalluWorld - Terminal: leaderboard
Metric: Hallucination rate (%) over the 529 HalluWorld-Terminal probes about recorded Terminal-Bench agent trajectories (golden answers grounded in file-system diffs, context up to 60k characters), temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (Medium) | 5.9 |
| 2 | O3 | 10.2 |
| 3 | O4 Mini | 14.2 |
| 4 | Claude Opus 4.6 (Medium) | 15.7 |
| 5 | Claude Sonnet 4.6 (Medium) | 17.8 |
| 6 | O3 Mini | 21 |
| 7 | GLM-5 (Non-reasoning) | 26.3 |
| 8 | GPT-4o | 38.4 |
| 9 | DeepSeek V3.1 (Non-reasoning) | 45.6 |
| 10 | GPT-5.4 Mini (Non-reasoning) | 50.3 |
| 11 | Qwen 3 30B A3B (Non-reasoning) | 51 |
| 12 | GPT-4o Mini | 51.8 |
| 13 | Kimi K2.6 (Non-reasoning) | 56.5 |
Interactive version: theaggregate.ai/benchmark?slug=halluworld-terminal · How It Works · Data refreshed daily, snapshot 2026-10-07.