POLAR-Bench - Privacy: leaderboard

Metric: Privacy score (%): share of the protected attributes the model kept hidden, trusted agent holding a source document, a privacy policy and a task, questioned single-turn or multi-turn by a fixed Llama-3.3-70B-Instruct third party under five attack strategies, attributes revealed in the transcript extracted by normalization and regex, mean over the 10 domains; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScore
1GPT-5.499.86
2Gemma 4 31B99.64
3Kimi K2.599.42
4GLM-5.199.23
5DeepSeek V3.199.17
6GLM-4.7 Flash98.9
7Llama 3.3 70B Instruct98.22
8Gemma 4 E4B95.45
9GPT-OSS-20B89
10GPT-OSS-120B88.19
11DeepSeek R1 Distill Qwen 32B84.79
12Gemma 3 27B81.61
13Qwen 3 32B77.3
14Gemma 4 E2B71.92
15Ministral 3 8B68.55

Interactive version: theaggregate.ai/benchmark?slug=polar-bench-privacy · How It Works · Data refreshed daily, snapshot 2026-10-07.