EvasionBench — leaderboard

Benchmark for evaluating model robustness against evasion-style safety and policy-circumvention prompts.

Metric: Macro-F1 (%). Source: iiiiqiiii.github.io. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1Gemini 3 Flash84.64
2Claude Opus 4.584.38
3GLM-4.782.91
4GPT-5.280.9
5Qwen3 Coder78.16
6MiniMax-M2.171.31
7DeepSeek V3.266.88
8Kimi K266.68

Interactive version: theaggregate.ai/benchmark?slug=evasionbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.