Llama 4 Maverick Instruct FP8: benchmark results

Llama 4 Maverick Instruct evaluated in FP8 quantization. Provider: Meta. Released 2025-04-05. Access: Open.

Unified ELO 1564 ± 1, rank #385 of 1392 rated models, from 126 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM SeaHELM - Flores (en-ta)51.98ChrF++100
HELM SeaHELM - Flores (th-en)59.27ChrF++100
HELM SeaHELM - XQuAD (Thai)77.14SQuAD macro-averaged F1 score100
HELM SeaHELM - IndicQA48.42SQuAD macro-averaged F1 score95
HELM Capabilities - IFEval90.76IFEval Strict Acc92
HELM SeaHELM - Flores (ta-en)57.09ChrF++90
HELM SeaHELM - MLHSD68.92Macro F1 score90
HELM SeaHELM - TyDiQA64.9SQuAD macro-averaged F1 score90
Kluster Hallucination Detection - RAG Method 2 Resistance96.66Resistance (100 - Hallucination Rate %)81.8
HELM Long Context - InfiniteBench En.MC89EM80
HELM Long Context - InfiniteBench En.Sum16.06ROUGE-L80
HELM Long Context - OpenAI MRCR21.49MRCR Accuracy80

Interactive version: theaggregate.ai/model?slug=llama-4-maverick-instruct-fp8 · How It Works · Data refreshed daily, snapshot 2026-09-05.