HelpBench - Moderation Tools: leaderboard

Metric: Moderation Tools topic (moderation and anti-abuse tools such as blocking, reporting and parental controls) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1GPT-5.384
2Gemini 3.1 Pro (Preview)82
3Gemini 3 Flash82
4Claude Opus 4.681
5GPT-5 Chat81
6Gemini 2.5 Pro80
7Claude Sonnet 4.680
8GLM-5.179
9GPT-4.178
10Claude Sonnet 478
11Qwen 3.6 Plus77
12Grok 476
13GLM-4.676
14DeepSeek V3.2 (Non-reasoning)74
15Grok 4.2073

Interactive version: theaggregate.ai/benchmark?slug=helpbench-moderation-tools · How It Works · Data refreshed daily, snapshot 2026-09-29.