PhantomBench - Terms (Meaning): leaderboard

Metric: Hallucination rate (%): share of answers that do not abstain on the Phantom-T subset (1,060 generated non-existent terms modeled on medical abbreviations, science glossaries and legal terms), verified absent from the 2.3-trillion-token Dolma corpus, when the prompt asks what the concept means or refers to (presupposing it exists); a Gemini 2.5 Flash judge labels abstention; reasoning-model rates count only generations that reached a final answer; lower is better. Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.

Top models

#ModelScore
1Qwen 3 14B8.02
2Qwen 3 4B8.87
3Qwen 3 8B9.06
4Qwen 2.5 7B Instruct10.94
5Qwen 3 32B13.68
6Qwen 3 1.7B22.76
7Llama 3.1 8B Instruct26.42
8Gemma 2 9B (IT)27.64
9Gemini 2.5 Flash33.21
10Gemini 2.5 Pro35
11GPT-OSS-20B (Medium)41.01
12Mistral 7B Instruct (v0.3)54.62
13GPT-OSS-20B (Low)54.77
14DeepSeek R1 Distill Qwen 32B66.96
15Gemma 3 12B (IT)87.26

Interactive version: theaggregate.ai/benchmark?slug=phantombench-terms-meaning · How It Works · Data refreshed daily, snapshot 2026-09-29.