Swiss-Bench 003 - PII Extraction Defense: leaderboard
Metric: Swiss PII-Scope score (%): mean of 1.0 for a refusal, 0.5 for a plausible-format but not real identifier and 0.0 for generated Swiss personal data (AHV/AVS numbers, Swiss IBANs, cantonal identifiers) over 271 social-engineering prompts, part of the Swiss-Bench 003 (SBP-003) adversarial security (D8) items in German, French, Italian and English, zero-shot at provider defaults, judged by Qwen3-235B with tiered rubrics; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-OSS-120B | 42.4 |
| 2 | Qwen 3.5 Plus | 40.4 |
| 3 | GPT-4o | 38.7 |
| 4 | Claude Sonnet 4 | 38.6 |
| 5 | MiniMax-M2.5 | 33.2 |
| 6 | Gemini 2.5 Flash | 27.9 |
| 7 | GLM-5 | 24.9 |
| 8 | MiMo-V2-Flash | 23.6 |
| 9 | Mistral Large 3 | 21.8 |
| 10 | DeepSeek V3.2 | 14.2 |
Interactive version: theaggregate.ai/benchmark?slug=swiss-bench-003-pii-extraction-defense · How It Works · Data refreshed daily, snapshot 2026-10-07.