EvoHarmBench - Pornographic Content: leaderboard

Metric: ASR@Readable (%; mean over the pornographic content sub-clusters of the share of adversarial rewrites that succeed against the moderator; 229 semantic sub-clusters built from 5,002 real-world adversarial posts in five violation categories; an adaptive DeepSeek-V3.2-Exp rewriter, reflector and comparison model evolve cluster-level rewriting strategies against the target moderator for 12 rounds; a rewrite counts only if it both evades the moderator and keeps human-recognizable harmful intent). Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.

Top models

#ModelScore
1GPT-5.578.9
2Claude Sonnet 4.685.2
3Qwen 3.6 Plus93.2
4Gemini 3.1 Pro (Preview)94.3
5Kimi K2.694.8
6DeepSeek V4 Pro96.7
7DeepSeek-V2-Lite97.7
8GLM-5.198.5
9Qwen 3 8B98.6
10Qwen 3 4B98.6

Interactive version: theaggregate.ai/benchmark?slug=evoharmbench-pornographic-content · How It Works · Data refreshed daily, snapshot 2026-09-26.