CodeGuard: leaderboard
Metric: F1 (%; three-way classification of 1,000 manually verified test prompts as irrelevant, relevant and safe, or relevant but unsafe; class definitions and one example per class in the prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.7 Sonnet | 64 |
| 2 | GPT-4o | 62 |
| 3 | Gemma 3 27B | 29 |
| 4 | Magistral Small | 20 |
Interactive version: theaggregate.ai/benchmark?slug=codeguard · How It Works · Data refreshed daily, snapshot 2026-09-26.