EvoHarmBench: leaderboard

Metric: ASR@Readable (%; sample-level share of adversarial rewrites that succeed against the moderator; 229 semantic sub-clusters built from 5,002 real-world adversarial posts in five violation categories; an adaptive DeepSeek-V3.2-Exp rewriter, reflector and comparison model evolve cluster-level rewriting strategies against the target moderator for 12 rounds; a rewrite counts only if it both evades the moderator and keeps human-recognizable harmful intent). Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)73
2GPT-5.573.1
3Claude Sonnet 4.673.2
4Qwen 3.6 Plus82.8
5Kimi K2.685.4
6DeepSeek V4 Pro86.2
7GLM-5.188.8
8DeepSeek-V2-Lite90.6
9Qwen 3 8B91.2
10Qwen 3 4B92.2

Interactive version: theaggregate.ai/benchmark?slug=evoharmbench · How It Works · Data refreshed daily, snapshot 2026-09-26.