Llama Guard 4 12B: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1461 ± 27, rank #927 of 1629 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ML-Bench (Multilingual Safety) - Attack-Enhanced Accuracy | 18 | Accuracy (%; share of refined unsafe queries with PAIR/AutoD | 59.1 |
| ML-Bench (Multilingual Safety) - Refined Query F1 | 27 | Unsafe-query F1 (%; safe/unsafe classification of refined qu | 54.5 |
| ATBench | 41.7 | Binary safe or unsafe classification F1 (%), unsafe as the p | 46.7 |
| SnakeBench | 18.2 | TrueSkill Rating | 32.4 |
| LM Market Cap LMC Score | 40 | LMC Score (0-100) | 28.4 |
| ML-Bench (Multilingual Safety) - Response F1 | 5 | Unsafe-response F1 (%; safe/unsafe classification of borderl | 27.3 |
| TRACE (LRM Safety) - Final Response | 62.92 | F1 on the unsafe class (%; guardrail safe or unsafe judgment | 23.5 |
| TRACE (LRM Safety) - Prompt | 66.52 | F1 on the unsafe class (%; guardrail safe or unsafe judgment | 23.5 |
| PolicyShiftBench - Adaptive Split | 29.2 | F1 of the block decision (%; policy-conditioned image guardr | 22.2 |
| PolicyShiftBench - Shift Split (PSS) | 1.1 | Policy Shift Score (%; over groups of the same image and ris | 13.9 |
| PolicyShiftBench | 19.6 | F1 of the block decision (%; policy-conditioned image guardr | 11.1 |
| PolicyShiftBench - Shift Split | 9.9 | F1 of the block decision (%; policy-conditioned image guardr | 11.1 |
Interactive version: theaggregate.ai/model?slug=llama-guard-4-12b · How It Works · Data refreshed daily, snapshot 2026-10-07.