Llama 3.1 405B Instruct: benchmark results

Meta Llama 3.1 405B instruction-tuned checkpoint. Provider: Meta. Released 2024-07-23. Access: Open.

Unified ELO 1558 ± 1, rank #406 of 1392 rated models, from 218 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM SeaHELM - XNLI (th)70EM100
HELM ThaiExam - TPAT-168.1EM100
HELM ThaiExam - A-Level66.93EM98.8
BenCzechMark83.19Average Score (%)98.6
HELM SeaHELM - IndicXNLI63.3EM97.5
MERA - ruDetox38.13Joint Score (%)97.1
MERA - RCB60.5Accuracy (%)96.2
MixEval66.2Score96.1
HELM ThaiExam - ThaiExam68.23EM95.1
HELM SeaHELM - Flores (en-ta)50.9ChrF++95
HELM SeaHELM - IndicSentiment98.3Macro F1 score95
HELM Lite88.91Mean win rate (self-reported)94.7

Interactive version: theaggregate.ai/model?slug=llama-3-1-405b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.