SSP-Bench - Hallucination: leaderboard

Metric: Non-hallucination rate (%; share of 1,281 generated Wikipedia-grounded factual questions answered correctly or explicitly declined, answers compared with the gold answer by an LLM judge; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.

Top models

#ModelScore
1GPT-5 Mini92.9
2Gemini 3 Flash (Preview)92.7
3Claude Haiku 4.591.8
4GPT-4o91
5Qwen 3.5 122B A10B91
6Grok 3 Mini91
7Qwen 3.5 35B A3B90.9
8Claude 3.5 Haiku90.9
9Llama 3.3 70B90.7
10Grok 4 Fast90.7
11Qwen 3.5 27B90
12DeepSeek R1 052889.5
13Grok 4.1 Fast89.3
14GPT-OSS-120B89.1
15Gemma 4 31B89.1

Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-hallucination · How It Works · Data refreshed daily, snapshot 2026-09-26.