HalluWorld - Grid Compound: leaderboard

Metric: Hallucination rate (%) on the cross-category compound probes of HalluWorld-Grid (MiniGrid worlds whose ground truth comes from the environment state), micro-averaged over levels and serializers, temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-5.5 (Medium)2.7
2Claude Opus 4.64.6
3Claude Opus 4.6 (Medium)4.7
4O35.7
5O4 Mini5.8
6Claude Sonnet 4.6 (Medium)7.5
7Claude Sonnet 4.67.6
8O3 Mini7.6
9Kimi K2.6 (Non-reasoning)8.8
10GPT-4o11.3
11DeepSeek V3 (0324)12.7
12GLM-5 (Non-reasoning)13.4
13Qwen 3 30B A3B (Non-reasoning)23
14GPT-5.4 Mini (Non-reasoning)24
15GPT-4o Mini28.7

Interactive version: theaggregate.ai/benchmark?slug=halluworld-grid-compound · How It Works · Data refreshed daily, snapshot 2026-10-07.