HelpBench - Harassment: leaderboard

Metric: Harassment topic (interpersonal harassment and abuse such as doxxing, stalking and impersonation) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1GPT-5 Chat88
2GPT-5.387
3GPT-4.185
4Claude Opus 4.685
5Grok 485
6Gemini 3 Flash85
7Qwen 3.6 Plus85
8Gemini 3.1 Pro (Preview)82
9Gemini 2.5 Pro82
10Claude Sonnet 4.682
11Claude Sonnet 482
12Grok 4.2082
13GLM-4.681
14GLM-5.180
15DeepSeek V3.2 (Non-reasoning)80

Interactive version: theaggregate.ai/benchmark?slug=helpbench-harassment · How It Works · Data refreshed daily, snapshot 2026-09-29.