PhantomBench - Entities (Date): leaderboard

Metric: Hallucination rate (%): share of answers that do not abstain on the Phantom-E subset (1,200 generated non-existent entities: events, holidays, elections, disasters, creative works and places), verified absent from the 2.3-trillion-token Dolma corpus, when the prompt asks when the concept originated or was introduced (presupposing it exists); a Gemini 2.5 Flash judge labels abstention; reasoning-model rates count only generations that reached a final answer; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct1.33
2Qwen 2.5 7B Instruct9.67
3Qwen 3 14B10
4Qwen 3 8B11.75
5Gemma 2 9B (IT)16.25
6Mistral 7B Instruct (v0.3)20.08
7Qwen 3 4B20.5
8Qwen 3 1.7B32.08
9Gemini 2.5 Pro33.25
10Qwen 3 32B33.42
11Gemini 2.5 Flash37.17
12DeepSeek R1 Distill Qwen 32B48.46
13GPT-OSS-20B (Low)56.18
14Gemma 3 12B (IT)77.33

Interactive version: theaggregate.ai/benchmark?slug=phantombench-entities-date · How It Works · Data refreshed daily, snapshot 2026-09-29.