HelpBench - Compromise: leaderboard

Metric: Compromise topic (compromised accounts and devices, malware and suspicious behavior) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1GPT-5.390
2GPT-5 Chat90
3GPT-4.187
4Claude Sonnet 486
5Qwen 3.6 Plus86
6Claude Sonnet 4.685
7Claude Opus 4.685
8Grok 484
9Gemini 3 Flash84
10GLM-5.183
11Grok 4.2083
12Gemini 3.1 Pro (Preview)82
13DeepSeek V3.2 (Non-reasoning)82
14Gemini 2.5 Pro81
15GLM-5-Turbo81

Interactive version: theaggregate.ai/benchmark?slug=helpbench-compromise · How It Works · Data refreshed daily, snapshot 2026-09-29.