Safety-Flag: leaderboard

Metric: Macro-F1 (%; mean over the seven source benchmarks of the macro-F1 of flag / do-not-flag verdicts on 198-200 class-balanced items from each of BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat and ToxiGen, recast as one flag / do-not-flag decision with harmonised source labels; one standardized two-option prompt with a JSON verdict and confidence; greedy decoding). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini85.3
2GPT-4.1 Mini84.3
3Gemma 2 9B (IT)82.4
4DeepSeek R1 Distill Llama 8B81.3
5Qwen 2.5 7B Instruct79.3
6Mistral 7B Instruct (v0.3)73.3
7OLMo-2-1124-7B-Instruct63.9
8Llama 3.1 8B Instruct44.3

Interactive version: theaggregate.ai/benchmark?slug=safety-flag · How It Works · Data refreshed daily, snapshot 2026-09-26.