HalluWorld - Grid Perceptual: leaderboard

Metric: Hallucination rate (%) on the perceptual probes (read-out of values present in the current observation, 6 levels) of HalluWorld-Grid (MiniGrid worlds whose ground truth comes from the environment state), micro-averaged over levels and serializers, temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1O30.3
2O3 Mini0.4
3GPT-5.5 (Medium)0.4
4O4 Mini0.8
5Claude Sonnet 4.6 (Medium)1.1
6Claude Opus 4.6 (Medium)1.7
7Claude Sonnet 4.62.8
8Claude Opus 4.63.7
9Kimi K2.6 (Non-reasoning)5.4
10GPT-4o10
11GPT-5.4 Mini (Non-reasoning)19.8
12GLM-5 (Non-reasoning)19.8
13Qwen 3 30B A3B (Non-reasoning)22.4
14DeepSeek V3 (0324)23.5
15GPT-4o Mini34.9

Interactive version: theaggregate.ai/benchmark?slug=halluworld-grid-perceptual · How It Works · Data refreshed daily, snapshot 2026-10-07.