HelpBench - Security Tools: leaderboard

Metric: Security Tools topic (security tools such as antivirus, firewalls, network security and sandboxing) rubric score (%): points for met positive criteria and avoided negative criteria (factual criteria written per question by privacy, safety and security experts, plus shared delivery criteria) over the maximum points, applied by a Gemini 2.5 Pro auto-rater at temperature 0; chat versions with default parameters, mean of five responses per question; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 18 models tracked.

Top models

#ModelScore
1Gemini 3 Flash86
2GPT-5 Chat85
3Gemini 3.1 Pro (Preview)84
4GPT-4.184
5GPT-5.384
6Claude Opus 4.683
7Gemini 2.5 Pro82
8Claude Sonnet 4.682
9Grok 482
10Claude Sonnet 482
11Qwen 3.6 Plus82
12GLM-5.182
13GLM-4.681
14DeepSeek V3.2 (Non-reasoning)81
15DeepSeek V3.1 (Non-reasoning)79

Interactive version: theaggregate.ai/benchmark?slug=helpbench-security-tools · How It Works · Data refreshed daily, snapshot 2026-09-29.