Llama Guard 4 12B: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1461 ± 27, rank #927 of 1629 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ML-Bench (Multilingual Safety) - Attack-Enhanced Accuracy18Accuracy (%; share of refined unsafe queries with PAIR/AutoD59.1
ML-Bench (Multilingual Safety) - Refined Query F127Unsafe-query F1 (%; safe/unsafe classification of refined qu54.5
ATBench41.7Binary safe or unsafe classification F1 (%), unsafe as the p46.7
SnakeBench18.2TrueSkill Rating32.4
LM Market Cap LMC Score40LMC Score (0-100)28.4
ML-Bench (Multilingual Safety) - Response F15Unsafe-response F1 (%; safe/unsafe classification of borderl27.3
TRACE (LRM Safety) - Final Response62.92F1 on the unsafe class (%; guardrail safe or unsafe judgment23.5
TRACE (LRM Safety) - Prompt66.52F1 on the unsafe class (%; guardrail safe or unsafe judgment23.5
PolicyShiftBench - Adaptive Split29.2F1 of the block decision (%; policy-conditioned image guardr22.2
PolicyShiftBench - Shift Split (PSS)1.1Policy Shift Score (%; over groups of the same image and ris13.9
PolicyShiftBench19.6F1 of the block decision (%; policy-conditioned image guardr11.1
PolicyShiftBench - Shift Split9.9F1 of the block decision (%; policy-conditioned image guardr11.1

Interactive version: theaggregate.ai/model?slug=llama-guard-4-12b · How It Works · Data refreshed daily, snapshot 2026-10-07.