Llama 4 Scout: benchmark results
Meta's open Llama 4 Scout, a 109B sparse MoE (17B active) long-context multimodal model (April 2025). Provider: Meta. Released 2025-04-05. Access: Open.
Unified ELO 1522 ± 1, rank #585 of 1392 rated models, from 318 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CARE-Bench | 50.4 | Overall (%) | 100 |
| AGC-Bench - sdat | 0.67 | Dataset z-score | 98.8 |
| AGC-Bench - unfun_corpus | 0.73 | Dataset z-score | 93.6 |
| Phare - Bias Resistance | 67.1 | Score (%) | 90.9 |
| AGC-Bench - c3_crosstalk | 1.07 | Dataset z-score | 89 |
| AGC-Bench - outline_to_story | 1.11 | Dataset z-score | 88.9 |
| AGC-Bench - creatset | 1.24 | Dataset z-score | 85.4 |
| AGC-Bench - scimon | 0.69 | Dataset z-score | 85.2 |
| LLM Stats (ChartQA) | 88.8 | Score (%) | 84 |
| MERA - ruDetox | 33.51 | Joint Score (%) | 83.7 |
| AndroidWorld | 91.4 | Success Rate pass@1 (%) | 80.8 |
| AGC-Bench - irfl | 0.69 | Dataset z-score | 78.1 |
Interactive version: theaggregate.ai/model?slug=llama-4-scout · How It Works · Data refreshed daily, snapshot 2026-09-05.