Llama 3.1 405B Instruct — benchmark results
Meta Llama 3.1 405B instruction-tuned checkpoint. Provider: Meta. Released 2024-07-23. Access: Open.
Unified ELO 1559 ± 12, rank #595 of 1776 rated models, from 205 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM SeaHELM - XNLI (th) | 70 | EM | 100 |
| HELM ThaiExam - TPAT-1 | 68.1 | EM | 100 |
| HELM ThaiExam - A-Level | 66.93 | EM | 98.8 |
| LiveBench Paraphrase | 81.13 | Score | 98.6 |
| LiveBench Simplify | 78.33 | Score | 98.6 |
| LiveBench Web Of Lies V2 | 80 | Score | 98.6 |
| BenCzechMark | 85.1 | Average Score (%) | 98.5 |
| HELM SeaHELM - IndicXNLI | 63.3 | EM | 97.5 |
| LiveBench Plot Unscrambling | 47.38 | Score | 97.2 |
| LiveBench Summarize | 78.93 | Score | 97.2 |
| MixEval | 66.2 | Score | 96.1 |
| LiveBench Cta | 60 | Score | 95.8 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-405b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.