Llama 4 Scout Instruct — benchmark results

Meta's smaller natively multimodal Llama 4 MoE (109B total/17B active, 16 experts) with a 10M-token context (April 2025). Provider: Meta. Released 2025-04-05. Access: Open.

Unified ELO 1550 ± 8, rank #619 of 1776 rated models, from 272 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Long Context - InfiniteBench En.Sum17.6ROUGE-L100
KOFFVQA - Hallucination and Robustness90Score (%)98.8
EuroEval Norwegian NLU - Norec63.43Sentiment classification Score (%)96.7
EuroEval Spanish NLU - Sentiment Headlines ES51.84Sentiment classification Score (%)95.4
EuroEval Swedish NLU - Swerec80Sentiment classification Score (%)95.4
FinBen - FinNum49.12Normalized Score95
FinBen - QA74.22Normalized Score94.7
EuroEval Swedish Common Sense Reasoning74.93Common Sense Reasoning Average Score (%)93.4
EuroEval Danish71.05Average Score (%)92.9
EuroEval Italian Common Sense Reasoning74.71Common Sense Reasoning Average Score (%)92.9
EuroEval Finnish NLU - Scandisent FI92.52Sentiment classification Score (%)92.6
EuroEval Spanish60.66Average Score (%)92.5

Interactive version: theaggregate.ai/model?slug=llama-4-scout-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.