SSP-Bench - Over-Refusal: leaderboard

Metric: Compliance rate (%; share of 684 generated benign boundary questions answered substantively rather than refused, judged by gpt-oss-safeguard-120b; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)100
2Gemma 3 27B99.7
3Gemma 3 12B99.6
4Gemma 3 4B99.6
5Grok 4 Fast99.6
6Grok 4.1 Fast99.4
7Llama 3.3 70B98.7
8Gemma 4 26B A4B98.5
9Grok 3 Mini98.4
10Gemma 4 31B98
11DeepSeek R1 052897.7
12Claude Haiku 4.596.8
13Gemma 3 1B96.3
14GPT-4o96.1
15GPT-5 Mini92.8

Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-over-refusal · How It Works · Data refreshed daily, snapshot 2026-09-26.