HelpBench - Platform Actions: leaderboard

Metric: Platform Actions topic (platform or third-party safety actions gone awry, such as bans and suspensions) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1GPT-5 Chat89
2GPT-5.386
3Claude Opus 4.685
4Claude Sonnet 4.684
5GPT-4.184
6Qwen 3.6 Plus84
7Grok 483
8Claude Sonnet 483
9Gemini 3.1 Pro (Preview)81
10Gemini 3 Flash81
11DeepSeek V3.2 (Non-reasoning)81
12GLM-5.180
13Gemini 2.5 Pro79
14Grok 4.2079
15DeepSeek V3.1 (Non-reasoning)79

Interactive version: theaggregate.ai/benchmark?slug=helpbench-platform-actions · How It Works · Data refreshed daily, snapshot 2026-09-29.