AgentHazard (Claude Code) - Data Exfiltration: leaderboard
Metric: Attack success rate (%) on the data exfiltration instances (unauthorized extraction of source code, environment variables, credentials or private data), with the model running inside the Claude Code CLI agent (persistent shell, file editing, tool use; restricted shell sessions), over AgentHazard instances (harmful objectives embedded in locally legitimate multi-step workflows, ten risk categories and ten attack strategies; the evaluated runs cover 1,000 instances), judged on the full trajectory (round_all) by Gemini-3-Flash, which labels a trajectory harmful when it realizes the objective; backbones served through an OpenAI-compatible API; lower is better. Source: arxiv.org. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 72B Instruct | 26.96 |
| 2 | Kimi K2 | 38.26 |
| 3 | Qwen 2.5 Coder 32B Instruct | 62.61 |
| 4 | Kimi K2.5 | 71.3 |
| 5 | Qwen 3 VL 235B A22B Instruct | 73.91 |
| 6 | GLM-4.6 | 85.22 |
Interactive version: theaggregate.ai/benchmark?slug=agenthazard-claude-code-data-exfiltration · How It Works · Data refreshed daily, snapshot 2026-10-07.