Llama 4 Scout Instruct: benchmark results

Meta's smaller natively multimodal Llama 4 MoE (109B total/17B active, 16 experts) with a 10M-token context (April 2025). Provider: Meta. Released 2025-04-05. Access: Open.

Unified ELO 1534 ± 1, rank #516 of 1392 rated models, from 311 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Long Context - InfiniteBench En.Sum17.6ROUGE-L100
KOFFVQA - Hallucination and Robustness90Score (%)98.8
EuroEval Norwegian NLU - Norec63.43Sentiment classification Score (%)96.7
EuroEval Spanish NLU - Sentiment Headlines ES51.84Sentiment classification Score (%)95.4
EuroEval Swedish NLU - Swerec80Sentiment classification Score (%)95.4
FinBen - FinNum49.12Normalized Score95
FinBen - QA74.22Normalized Score94.7
EuroEval Swedish Common Sense Reasoning74.93Common Sense Reasoning Average Score (%)93.4
EuroEval Danish71.05Average Score (%)92.9
EuroEval Italian Common Sense Reasoning74.71Common Sense Reasoning Average Score (%)92.9
EuroEval Finnish NLU - Scandisent FI92.52Sentiment classification Score (%)92.6
EuroEval Spanish60.66Average Score (%)92.5

Interactive version: theaggregate.ai/model?slug=llama-4-scout-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.