PhantomBench - Terms (Place): leaderboard

Metric: Hallucination rate (%): share of answers that do not abstain on the Phantom-T subset (1,060 generated non-existent terms modeled on medical abbreviations, science glossaries and legal terms), verified absent from the 2.3-trillion-token Dolma corpus, when the prompt asks where the concept was discovered or took place (presupposing it exists); a Gemini 2.5 Flash judge labels abstention; reasoning-model rates count only generations that reached a final answer; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Qwen 3 8B3.21
2Qwen 3 14B3.3
3Llama 3.1 8B Instruct3.58
4Qwen 3 4B6.23
5Qwen 2.5 7B Instruct6.51
6Qwen 3 1.7B9.15
7Gemma 2 9B (IT)12.83
8Qwen 3 32B14.06
9Gemini 2.5 Pro15.75
10Gemini 2.5 Flash19.34
11Mistral 7B Instruct (v0.3)43.68
12DeepSeek R1 Distill Qwen 32B58.97
13GPT-OSS-20B (Low)72.68
14Gemma 3 12B (IT)85.28

Interactive version: theaggregate.ai/benchmark?slug=phantombench-terms-place · How It Works · Data refreshed daily, snapshot 2026-09-29.