GMP - Co-occurring Violations (Micro-F1): leaderboard

Metric: Micro-F1 (x100) over the 12 violation categories of GMP Task A (identify every co-occurring violation in a post; 1,400 posts, 980 unsafe, 81% of them multi-label), zero-shot prompt with JSON output; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 483.47#194
2GLM-4.5 (Non-reasoning)83#265 (GLM-4.5)
3DeepSeek V3.1 (Non-reasoning)82.58#260 (DeepSeek V3.1)
4Qwen 2.5 VL 72B Instruct78.21#364
5GLM-4.5 (Thinking)78#265 (GLM-4.5)
6Gemini 2.5 Flash77.43#237
7GPT-4o Mini76.97#588
8Kimi K272#236
9Qwen 3 30B A3B 2507 Instruct70.17#464
10DeepSeek V3.1 (Thinking)68#260 (DeepSeek V3.1)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gmp-co-occurring-violations-micro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.