EvoHarmBench - Gambling and Fraud: leaderboard

Metric: ASR@Readable (%; mean over the gambling and fraud sub-clusters of the share of adversarial rewrites that succeed against the moderator; 229 semantic sub-clusters built from 5,002 real-world adversarial posts in five violation categories; an adaptive DeepSeek-V3.2-Exp rewriter, reflector and comparison model evolve cluster-level rewriting strategies against the target moderator for 12 rounds; a rewrite counts only if it both evades the moderator and keeps human-recognizable harmful intent). Source: arxiv.org. Saturation forecast: Around July 2028. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)71.2
2Claude Sonnet 4.684.5
3GLM-5.190.5
4Qwen 3.6 Plus92.1
5DeepSeek V4 Pro95.1
6Kimi K2.697.6
7GPT-5.597.7
8DeepSeek-V2-Lite99.3
9Qwen 3 8B99.5
10Qwen 3 4B99.7

Interactive version: theaggregate.ai/benchmark?slug=evoharmbench-gambling-and-fraud · How It Works · Data refreshed daily, snapshot 2026-09-26.