SSP-Bench - Privacy: leaderboard
Metric: Safe rate (%; share of 5,000 PII-extraction prompts over legal and medical documents filled with synthetic PII, six attack framings, on which no targeted PII span appears verbatim in the response; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Around 2032. 24 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 53.9 |
| 2 | GPT-5 Mini | 52.5 |
| 3 | Claude 3.5 Haiku | 52 |
| 4 | GPT-4o | 47 |
| 5 | Grok 4 Fast | 45.5 |
| 6 | Gemma 4 31B | 45.3 |
| 7 | Phi-4 | 43.8 |
| 8 | Gemma 4 26B A4B | 41.4 |
| 9 | Gemini 3 Flash (Preview) | 40.9 |
| 10 | Grok 4.1 Fast | 40.2 |
| 11 | Gemma 3 12B | 37.5 |
| 12 | Llama 3.3 70B | 35.3 |
| 13 | Grok 3 Mini | 34.5 |
| 14 | GPT-OSS-120B | 33.1 |
| 15 | GPT-OSS-20B | 33.1 |
Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-privacy · How It Works · Data refreshed daily, snapshot 2026-09-26.