Llama 2 70B: benchmark results
Provider: Meta. Released 2023-07-18. Access: Open.
Unified ELO 1439 ± 1, rank #993 of 1392 rated models, from 54 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM | 94.35 | Mean win rate (self-reported) | 100 |
| HELM Classic - NarrativeQA | 76.99 | F1 (%) | 100 |
| HELM Classic - NaturalQuestions Closed Book | 45.84 | F1 (%) | 100 |
| HELM Classic - WikiFact | 44.92 | Exact Match (%) | 100 |
| HELM Classic - bAbI | 70.53 | Exact Match (%) | 100 |
| LLM Benchmarker Suite | 62.53 | Average Score (%) | 100 |
| HELM Classic - BoolQ | 88.6 | Exact Match (%) | 98.5 |
| ToolBench - VirtualHome | 24.74 | Task Score | 97.7 |
| HELM Classic - IMDB | 96.1 | Exact Match (%) | 95.5 |
| HELM Classic - MMLU | 58.17 | Exact Match (%) | 95.5 |
| HELM Classic - QuAC | 48.44 | F1 (%) | 95.4 |
| ToolBench - WebShop Long | 1.53 | Task Score | 95.3 |
Interactive version: theaggregate.ai/model?slug=llama-2-70b · How It Works · Data refreshed daily, snapshot 2026-09-05.