EvoHarmBench - Spam and Flooding: leaderboard

Metric: ASR@Readable (%; mean over the spam and flooding sub-clusters of the share of adversarial rewrites that succeed against the moderator; 229 semantic sub-clusters built from 5,002 real-world adversarial posts in five violation categories; an adaptive DeepSeek-V3.2-Exp rewriter, reflector and comparison model evolve cluster-level rewriting strategies against the target moderator for 12 rounds; a rewrite counts only if it both evades the moderator and keeps human-recognizable harmful intent). Source: arxiv.org. Saturation forecast: Around 2036. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.661.7
2GPT-5.565.8
3DeepSeek-V2-Lite72.3
4Gemini 3.1 Pro (Preview)72.4
5Kimi K2.681.8
6Qwen 3 4B83.6
7Qwen 3 8B85
8Qwen 3.6 Plus86
9DeepSeek V4 Pro86.8
10GLM-5.191.6

Interactive version: theaggregate.ai/benchmark?slug=evoharmbench-spam-and-flooding · How It Works · Data refreshed daily, snapshot 2026-09-26.