PhantomBench - Entities (Place): leaderboard

Metric: Hallucination rate (%): share of answers that do not abstain on the Phantom-E subset (1,200 generated non-existent entities: events, holidays, elections, disasters, creative works and places), verified absent from the 2.3-trillion-token Dolma corpus, when the prompt asks where the concept was discovered or took place (presupposing it exists); a Gemini 2.5 Flash judge labels abstention; reasoning-model rates count only generations that reached a final answer; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct3.67
2Qwen 2.5 7B Instruct12.25
3Qwen 3 14B16.08
4Qwen 3 8B18.59
5Gemma 2 9B (IT)24.42
6Qwen 3 4B24.84
7Qwen 3 1.7B33.25
8Mistral 7B Instruct (v0.3)35.16
9Gemini 2.5 Pro36.17
10Gemini 2.5 Flash40.34
11Qwen 3 32B43.09
12DeepSeek R1 Distill Qwen 32B63.95
13GPT-OSS-20B (Medium)67.53
14GPT-OSS-20B (Low)80.39
15Gemma 3 12B (IT)85.17

Interactive version: theaggregate.ai/benchmark?slug=phantombench-entities-place · How It Works · Data refreshed daily, snapshot 2026-09-29.