HalluWorld - Terminal: leaderboard

Metric: Hallucination rate (%) over the 529 HalluWorld-Terminal probes about recorded Terminal-Bench agent trajectories (golden answers grounded in file-system diffs, context up to 60k characters), temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1GPT-5.5 (Medium)5.9
2O310.2
3O4 Mini14.2
4Claude Opus 4.6 (Medium)15.7
5Claude Sonnet 4.6 (Medium)17.8
6O3 Mini21
7GLM-5 (Non-reasoning)26.3
8GPT-4o38.4
9DeepSeek V3.1 (Non-reasoning)45.6
10GPT-5.4 Mini (Non-reasoning)50.3
11Qwen 3 30B A3B (Non-reasoning)51
12GPT-4o Mini51.8
13Kimi K2.6 (Non-reasoning)56.5

Interactive version: theaggregate.ai/benchmark?slug=halluworld-terminal · How It Works · Data refreshed daily, snapshot 2026-10-07.