PhantomBench - Terms (Date): leaderboard

Metric: Hallucination rate (%): share of answers that do not abstain on the Phantom-T subset (1,060 generated non-existent terms modeled on medical abbreviations, science glossaries and legal terms), verified absent from the 2.3-trillion-token Dolma corpus, when the prompt asks when the concept originated or was introduced (presupposing it exists); a Gemini 2.5 Flash judge labels abstention; reasoning-model rates count only generations that reached a final answer; lower is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct0.94
2Qwen 3 14B1.79
3Gemini 2.5 Pro2.26
4Qwen 3 8B3.87
5Gemma 2 9B (IT)4.53
6Qwen 2.5 7B Instruct5.57
7Qwen 3 32B6.89
8Qwen 3 4B9.53
9Qwen 3 1.7B12.64
10Gemini 2.5 Flash20.85
11GPT-OSS-20B (Medium)30.5
12Mistral 7B Instruct (v0.3)32.08
13DeepSeek R1 Distill Qwen 32B43.08
14GPT-OSS-20B (Low)55.01
15Gemma 3 12B (IT)69.25

Interactive version: theaggregate.ai/benchmark?slug=phantombench-terms-date · How It Works · Data refreshed daily, snapshot 2026-09-29.