SSP-Bench - Privacy: leaderboard

Metric: Safe rate (%; share of 5,000 PII-extraction prompts over legal and medical documents filled with synthetic PII, six attack framings, on which no targeted PII span appears verbatim in the response; SSP-Bench dynamic benchmark instance scored on the 24-model testing panel, held out from item generation and selection; proprietary models through OpenRouter, March-April 2026). Source: arxiv.org. Saturation forecast: Around 2032. 24 models tracked.

Top models

#ModelScore
1Claude Haiku 4.553.9
2GPT-5 Mini52.5
3Claude 3.5 Haiku52
4GPT-4o47
5Grok 4 Fast45.5
6Gemma 4 31B45.3
7Phi-443.8
8Gemma 4 26B A4B41.4
9Gemini 3 Flash (Preview)40.9
10Grok 4.1 Fast40.2
11Gemma 3 12B37.5
12Llama 3.3 70B35.3
13Grok 3 Mini34.5
14GPT-OSS-120B33.1
15GPT-OSS-20B33.1

Interactive version: theaggregate.ai/benchmark?slug=ssp-bench-privacy · How It Works · Data refreshed daily, snapshot 2026-09-26.