GuardZoo (Zero-Shot) - Safe/Unsafe Accuracy: leaderboard
Metric: Binary safe/unsafe accuracy (%; GuardZoo in-domain test split, 20% of the Aegis 2.0, WildGuardMix, BeaverTails, S-Eval, PolicyGuard-Bench, ASSE-Bench and R-Judge samples, relabeled; zero-shot guard models). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3Guard-4B | 77 |
| 2 | Qwen3Guard-8B | 76.14 |
| 3 | Llama Guard 3 8B | 70 |
| 4 | ShieldGemma 2B | 68.09 |
| 5 | LlamaGuard 2 8B | 67.71 |
| 6 | Qwen3Guard-0.6B | 66.5 |
| 7 | ShieldGemma 9B | 63.74 |
| 8 | ShieldGemma 27B | 52.07 |
| 9 | Llama Guard 3 1B | 45.13 |
Interactive version: theaggregate.ai/benchmark?slug=guardzoo-zero-shot-safe-unsafe-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-26.