SSP-Bench - Hallucination: leaderboard
Metric: Non-hallucination rate (%; share of 1,281 generated Wikipedia-grounded factual questions answered correctly or explicitly declined, answers compared with the gold answer by an LLM judge; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 92.9 |
| 2 | Gemini 3 Flash (Preview) | 92.7 |
| 3 | Claude Haiku 4.5 | 91.8 |
| 4 | GPT-4o | 91 |
| 5 | Qwen 3.5 122B A10B | 91 |
| 6 | Grok 3 Mini | 91 |
| 7 | Qwen 3.5 35B A3B | 90.9 |
| 8 | Claude 3.5 Haiku | 90.9 |
| 9 | Llama 3.3 70B | 90.7 |
| 10 | Grok 4 Fast | 90.7 |
| 11 | Qwen 3.5 27B | 90 |
| 12 | DeepSeek R1 0528 | 89.5 |
| 13 | Grok 4.1 Fast | 89.3 |
| 14 | GPT-OSS-120B | 89.1 |
| 15 | Gemma 4 31B | 89.1 |
Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-hallucination · How It Works · Data refreshed daily, snapshot 2026-09-26.