HalluWorld - Chess (Incorrect FEN): leaderboard

Metric: Hallucination rate (%) over the seven HalluWorld-Chess probes (about 348 questions) when the prompt also includes an incorrect FEN string, temperature 0 where supported, thinking models at medium effort with up to 16k thinking tokens, every model at most 256 answer tokens; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1GPT-5.44
2GPT-5.5 (Medium)4
3O36
4O4 Mini8.3
5GPT-5.4 Mini (Medium)22.1
6O3 Mini24.7
7Claude Opus 4.6 (Medium)25.9
8Claude Sonnet 4.6 (Medium)26.1
9Claude Opus 4.633
10Claude Sonnet 4.639.7
11Kimi K2.6 (Non-reasoning)42.2
12GLM-5 (Non-reasoning)47.4
13GPT-5.5 (Non-reasoning)48.9
14DeepSeek V3.1 (Non-reasoning)50
15GPT-5.4 Mini (Non-reasoning)54.3

Interactive version: theaggregate.ai/benchmark?slug=halluworld-chess-incorrect-fen · How It Works · Data refreshed daily, snapshot 2026-10-07.