HarmActionsEval — leaderboard

Agent safety benchmark measuring how often autonomous LLM agents avoid harmful tool actions under adversarial pressure.

Metric: SafeActions@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 10 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B23.4
2Claude Sonnet 4.62.84
3Phi-4-mini (Reasoning)2.84
4Ministral 3 3B2.13
5Gemini 3.1 Flash Lite0.71
6GPT-5.4 Mini0.71
7Claude Haiku 4.50
8Phi-4 Mini Instruct0

Interactive version: theaggregate.ai/benchmark?slug=harmactionseval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.