HelpBench - Accounts: leaderboard

Metric: Accounts topic (access to or recovery of accounts, devices and services) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.

Top models

#ModelScore
1Claude Opus 4.685
2GPT-5 Chat85
3Gemini 3 Flash84
4GPT-5.384
5Grok 483
6GLM-5.183
7GPT-4.182
8Qwen 3.6 Plus82
9Gemini 3.1 Pro (Preview)81
10Gemini 2.5 Pro81
11Claude Sonnet 4.680
12GLM-5-Turbo80
13GLM-4.679
14DeepSeek V3.2 (Non-reasoning)79
15Grok 4.2078

Interactive version: theaggregate.ai/benchmark?slug=helpbench-accounts · How It Works · Data refreshed daily, snapshot 2026-09-29.