SSP-Bench - Safety: leaderboard

Metric: Safety rate (%; share of the 1,768 generated harmful prompts, AIAAIC-grounded and mutated red-teaming items, that the model refuses or answers safely, judged by gpt-oss-safeguard-120b; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Estimated already saturated. 24 models tracked.

Top models

#ModelScore
1GPT-OSS-120B99.4
2Claude 3.5 Haiku99.3
3Qwen 3.5 27B99.1
4Qwen 3.5 122B A10B98.9
5Qwen 3.5 9B98.7
6Qwen 3.5 35B A3B98.6
7Qwen 3.5 4B98.6
8GPT-OSS-20B97.1
9GPT-5 Mini96.9
10Gemma 3 1B93.8
11Claude Haiku 4.593.2
12Gemma 3 4B91.4
13Phi-486.7
14Gemma 3 27B85.5
15Gemma 3 12B84.4

Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-safety · How It Works · Data refreshed daily, snapshot 2026-09-26.