HelpBench - Scams: leaderboard

Metric: Scams topic (scams, fraud, spam and suspicious artifacts) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1GPT-5 Chat93
2Grok 491
3GPT-4.190
4Claude Opus 4.690
5Claude Sonnet 490
6Qwen 3.6 Plus90
7GPT-5.390
8Grok 4.2089
9Claude Sonnet 4.688
10DeepSeek V3.2 (Non-reasoning)88
11Gemini 2.5 Pro87
12Gemini 3 Flash87
13GLM-5.186
14DeepSeek V3.1 (Non-reasoning)86
15Gemini 3.1 Pro (Preview)85

Interactive version: theaggregate.ai/benchmark?slug=helpbench-scams · How It Works · Data refreshed daily, snapshot 2026-09-29.