HalluWorld - Grid Uncertainty: leaderboard

Metric: Hallucination rate (%) on the uncertainty probes (abstaining when the evidence is insufficient, 5 levels) of HalluWorld-Grid (MiniGrid worlds whose ground truth comes from the environment state), micro-averaged over levels and serializers, temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GLM-5 (Non-reasoning)0
2GPT-5.5 (Medium)0.1
3O30.3
4GPT-4o0.5
5O4 Mini0.9
6Claude Sonnet 4.61.1
7Claude Sonnet 4.6 (Medium)1.1
8Kimi K2.6 (Non-reasoning)1.9
9Claude Opus 4.63.4
10Claude Opus 4.6 (Medium)3.4
11DeepSeek V3 (0324)4.6
12GPT-4o Mini5.1
13GPT-5.4 Mini (Non-reasoning)5.7
14O3 Mini5.9
15Qwen 3 30B A3B (Non-reasoning)13.2

Interactive version: theaggregate.ai/benchmark?slug=halluworld-grid-uncertainty · How It Works · Data refreshed daily, snapshot 2026-10-07.