Llama 4 Maverick Instruct FP8 — benchmark results
Llama 4 Maverick Instruct evaluated in FP8 quantization. Provider: Meta. Released 2025-04-05. Access: Open.
Unified ELO 1520 ± 24, rank #724 of 1776 rated models, from 98 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM SeaHELM - Flores (en-ta) | 51.98 | ChrF++ | 100 |
| HELM SeaHELM - Flores (th-en) | 59.27 | ChrF++ | 100 |
| HELM SeaHELM - XQuAD (Thai) | 77.14 | SQuAD macro-averaged F1 score | 100 |
| HELM SeaHELM - IndicQA | 48.42 | SQuAD macro-averaged F1 score | 95 |
| HELM Capabilities - IFEval | 90.76 | IFEval Strict Acc | 92 |
| HELM SeaHELM - Flores (ta-en) | 57.09 | ChrF++ | 90 |
| HELM SeaHELM - MLHSD | 68.92 | Macro F1 score | 90 |
| HELM SeaHELM - TyDiQA | 64.9 | SQuAD macro-averaged F1 score | 90 |
| CPTU Bench | 3.93 | Average Score (1-5) | 84.2 |
| Kluster Hallucination Detection - RAG Method 2 Resistance | 96.66 | Resistance (100 - Hallucination Rate %) | 81.8 |
| HELM Long Context - InfiniteBench En.MC | 89 | EM | 80 |
| HELM Long Context - InfiniteBench En.Sum | 16.06 | ROUGE-L | 80 |
Interactive version: theaggregate.ai/model?slug=llama-4-maverick-instruct-fp8 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.