Llama 4 Scout Instruct — benchmark results
Meta's smaller natively multimodal Llama 4 MoE (109B total/17B active, 16 experts) with a 10M-token context (April 2025). Provider: Meta. Released 2025-04-05. Access: Open.
Unified ELO 1550 ± 8, rank #619 of 1776 rated models, from 272 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Long Context - InfiniteBench En.Sum | 17.6 | ROUGE-L | 100 |
| KOFFVQA - Hallucination and Robustness | 90 | Score (%) | 98.8 |
| EuroEval Norwegian NLU - Norec | 63.43 | Sentiment classification Score (%) | 96.7 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 51.84 | Sentiment classification Score (%) | 95.4 |
| EuroEval Swedish NLU - Swerec | 80 | Sentiment classification Score (%) | 95.4 |
| FinBen - FinNum | 49.12 | Normalized Score | 95 |
| FinBen - QA | 74.22 | Normalized Score | 94.7 |
| EuroEval Swedish Common Sense Reasoning | 74.93 | Common Sense Reasoning Average Score (%) | 93.4 |
| EuroEval Danish | 71.05 | Average Score (%) | 92.9 |
| EuroEval Italian Common Sense Reasoning | 74.71 | Common Sense Reasoning Average Score (%) | 92.9 |
| EuroEval Finnish NLU - Scandisent FI | 92.52 | Sentiment classification Score (%) | 92.6 |
| EuroEval Spanish | 60.66 | Average Score (%) | 92.5 |
Interactive version: theaggregate.ai/model?slug=llama-4-scout-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.