POLAR-Bench - Utility: leaderboard

Metric: Utility score (%): share of the task-relevant attributes the model shared to complete the task, trusted agent holding a source document, a privacy policy and a task, questioned single-turn or multi-turn by a fixed Llama-3.3-70B-Instruct third party under five attack strategies, attributes revealed in the transcript extracted by normalization and regex, mean over the 10 domains; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 22 models tracked.

Top models

#ModelScore
1GLM-5.189.46
2Gemma 4 E2B87.56
3Gemma 4 E4B86.29
4GPT-5.484.79
5Gemma 3 27B84.71
6Ministral 3 14B84.48
7Qwen 3 32B84.43
8Ministral 3 8B83.7
9DeepSeek R1 Distill Llama 70B82.93
10Gemma 4 31B82.77
11DeepSeek V3.182.44
12Apertus-70B-Instruct-250981.8
13Ministral 3 3B81.46
14SmolLM3-3B77.66
15DeepSeek R1 Distill Qwen 32B77.6

Interactive version: theaggregate.ai/benchmark?slug=polar-bench-utility · How It Works · Data refreshed daily, snapshot 2026-10-07.