CodeGuard: leaderboard

Metric: F1 (%; three-way classification of 1,000 manually verified test prompts as irrelevant, relevant and safe, or relevant but unsafe; class definitions and one example per class in the prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet64
2GPT-4o62
3Gemma 3 27B29
4Magistral Small20

Interactive version: theaggregate.ai/benchmark?slug=codeguard · How It Works · Data refreshed daily, snapshot 2026-09-26.