SSP-Bench - Over-Refusal: leaderboard
Metric: Compliance rate (%; share of 684 generated benign boundary questions answered substantively rather than refused, judged by gpt-oss-safeguard-120b; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 100 |
| 2 | Gemma 3 27B | 99.7 |
| 3 | Gemma 3 12B | 99.6 |
| 4 | Gemma 3 4B | 99.6 |
| 5 | Grok 4 Fast | 99.6 |
| 6 | Grok 4.1 Fast | 99.4 |
| 7 | Llama 3.3 70B | 98.7 |
| 8 | Gemma 4 26B A4B | 98.5 |
| 9 | Grok 3 Mini | 98.4 |
| 10 | Gemma 4 31B | 98 |
| 11 | DeepSeek R1 0528 | 97.7 |
| 12 | Claude Haiku 4.5 | 96.8 |
| 13 | Gemma 3 1B | 96.3 |
| 14 | GPT-4o | 96.1 |
| 15 | GPT-5 Mini | 92.8 |
Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-over-refusal · How It Works · Data refreshed daily, snapshot 2026-09-26.