EvoHarmBench - Abusive Content: leaderboard

Metric: ASR@Readable (%; mean over the abusive content sub-clusters of the share of adversarial rewrites that succeed against the moderator; 229 semantic sub-clusters built from 5,002 real-world adversarial posts in five violation categories; an adaptive DeepSeek-V3.2-Exp rewriter, reflector and comparison model evolve cluster-level rewriting strategies against the target moderator for 12 rounds; a rewrite counts only if it both evades the moderator and keeps human-recognizable harmful intent). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1GPT-5.546.8
2Gemini 3.1 Pro (Preview)48.3
3Claude Sonnet 4.657.4
4Qwen 3.6 Plus62.2
5DeepSeek V4 Pro71.1
6Kimi K2.673.1
7GLM-5.178.6
8Qwen 3 8B93.2
9DeepSeek-V2-Lite96.8
10Qwen 3 4B98.3

Interactive version: theaggregate.ai/benchmark?slug=evoharmbench-abusive-content · How It Works · Data refreshed daily, snapshot 2026-09-26.