Llama 4 Maverick Instruct FP8 — benchmark results

Llama 4 Maverick Instruct evaluated in FP8 quantization. Provider: Meta. Released 2025-04-05. Access: Open.

Unified ELO 1520 ± 24, rank #724 of 1776 rated models, from 98 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM SeaHELM - Flores (en-ta)51.98ChrF++100
HELM SeaHELM - Flores (th-en)59.27ChrF++100
HELM SeaHELM - XQuAD (Thai)77.14SQuAD macro-averaged F1 score100
HELM SeaHELM - IndicQA48.42SQuAD macro-averaged F1 score95
HELM Capabilities - IFEval90.76IFEval Strict Acc92
HELM SeaHELM - Flores (ta-en)57.09ChrF++90
HELM SeaHELM - MLHSD68.92Macro F1 score90
HELM SeaHELM - TyDiQA64.9SQuAD macro-averaged F1 score90
CPTU Bench3.93Average Score (1-5)84.2
Kluster Hallucination Detection - RAG Method 2 Resistance96.66Resistance (100 - Hallucination Rate %)81.8
HELM Long Context - InfiniteBench En.MC89EM80
HELM Long Context - InfiniteBench En.Sum16.06ROUGE-L80

Interactive version: theaggregate.ai/model?slug=llama-4-maverick-instruct-fp8 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.