GMP - Co-occurring Violations: leaderboard

Metric: Macro-F1 (x100) over the 12 violation categories of GMP Task A (identify every co-occurring violation in a post; 1,400 posts, 980 unsafe, 81% of them multi-label), zero-shot prompt with JSON output; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.

Top models

#ModelScoreOverall rank
1GLM-4.5 (Non-reasoning)58#265 (GLM-4.5)
2Claude Sonnet 457.58#194
3DeepSeek V3.1 (Non-reasoning)57.37#260 (DeepSeek V3.1)
4Qwen 2.5 VL 72B Instruct54.94#364
5GPT-4o Mini53.72#588
6GLM-4.5 (Thinking)52#265 (GLM-4.5)
7Gemini 2.5 Flash51.87#237
8Qwen 3 30B A3B 2507 Instruct50.71#464
9Kimi K249#236
10DeepSeek V3.1 (Thinking)47#260 (DeepSeek V3.1)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gmp-co-occurring-violations · How It Works · Data refreshed daily, snapshot 2026-10-11.