PhantomBench - Terms (Existence): leaderboard

Metric: Hallucination rate (%): share of answers that do not abstain on the Phantom-T subset (1,060 generated non-existent terms modeled on medical abbreviations, science glossaries and legal terms), verified absent from the 2.3-trillion-token Dolma corpus, when the prompt asks whether the concept exists; a Gemini 2.5 Flash judge labels abstention; reasoning-model rates count only generations that reached a final answer; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-OSS-20B (Medium)2.42
2Gemini 2.5 Pro2.83
3Qwen 3 14B3.21
4Qwen 3 32B3.77
5Qwen 3 4B4.06
6GPT-OSS-20B (Low)4.06
7Qwen 3 8B4.91
8Qwen 2.5 7B Instruct5.57
9Gemma 2 9B (IT)7.45
10DeepSeek R1 Distill Qwen 32B9.35
11Llama 3.1 8B Instruct14.34
12Qwen 3 1.7B16.79
13Gemini 2.5 Flash17.74
14Mistral 7B Instruct (v0.3)30.47
15Gemma 3 12B (IT)33.3

Interactive version: theaggregate.ai/benchmark?slug=phantombench-terms-existence · How It Works · Data refreshed daily, snapshot 2026-09-29.