EvoHarmBench (Direct Audit): leaderboard

Metric: Miss rate (%; share of the 5,002 real-world harmful posts in five violation categories that the moderator fails to flag when they are submitted unchanged, with the same category-specific moderation prompts as the iterative protocol). Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScore
1GPT-5.531.7
2Claude Sonnet 4.636.2
3Gemini 3.1 Pro (Preview)37.1
4Qwen 3.6 Plus39.5
5Kimi K2.639.6
6DeepSeek V4 Pro41.3
7GLM-5.143.3

Interactive version: theaggregate.ai/benchmark?slug=evoharmbench-direct-audit · How It Works · Data refreshed daily, snapshot 2026-09-26.