HalluWorld - Grid Memory: leaderboard

Metric: Hallucination rate (%) on the memory probes (tracking states across earlier observations, 6 levels) of HalluWorld-Grid (MiniGrid worlds whose ground truth comes from the environment state), micro-averaged over levels and serializers, temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.63.8
2Claude Opus 4.6 (Medium)3.8
3O4 Mini4.5
4O3 Mini4.9
5Claude Sonnet 4.65
6O35.5
7GPT-5.5 (Medium)5.9
8Kimi K2.6 (Non-reasoning)6.2
9Claude Sonnet 4.6 (Medium)6.2
10DeepSeek V3 (0324)14.3
11GLM-5 (Non-reasoning)15.8
12GPT-4o17.2
13GPT-5.4 Mini (Non-reasoning)26.6
14Qwen 3 30B A3B (Non-reasoning)27.7
15GPT-4o Mini39

Interactive version: theaggregate.ai/benchmark?slug=halluworld-grid-memory · How It Works · Data refreshed daily, snapshot 2026-10-07.