HalluWorld - Grid Causal: leaderboard

Metric: Hallucination rate (%) on the causal probes (forward simulation of fire, flood and pressure-plate mechanics, 9 levels) of HalluWorld-Grid (MiniGrid worlds whose ground truth comes from the environment state), micro-averaged over levels and serializers, temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Around August 2028. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.619.7
2Claude Opus 4.620.1
3GPT-5.4 Mini (Non-reasoning)22.2
4O3 Mini23.8
5Claude Sonnet 4.6 (Medium)24
6Claude Opus 4.6 (Medium)25.2
7GPT-5.5 (Medium)25.3
8O326
9GLM-5 (Non-reasoning)26.4
10GPT-4o29
11O4 Mini30.6
12GPT-4o Mini30.9
13Qwen 3 30B A3B (Non-reasoning)31.6
14DeepSeek V3 (0324)32.2
15Kimi K2.6 (Non-reasoning)47.4

Interactive version: theaggregate.ai/benchmark?slug=halluworld-grid-causal · How It Works · Data refreshed daily, snapshot 2026-10-07.