Llama 2 13B — benchmark results
Provider: Meta. Released 2023-07-18. Access: Open.
Unified ELO 1329 ± 14, rank #1473 of 1776 rated models, from 72 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - IMDB | 96.2 | Exact Match (%) | 98.5 |
| HELM Classic - Entity Data Imputation | 84.36 | Exact Match (%) | 97.7 |
| ToolBench - WebShop Short | 31.67 | Task Score | 97.7 |
| HELM Classic - NarrativeQA | 74.4 | F1 (%) | 93.8 |
| HELM Classic - LSAT | 23.48 | Exact Match (%) | 92.6 |
| HELM Classic - WikiFact | 36.47 | Exact Match (%) | 90.9 |
| ToolBench - WebShop Long | 0.6 | Task Score | 90.7 |
| HELM | 82.3 | Mean win rate (self-reported) | 89.1 |
| OpenEval - XSum | 30.43 | ROUGE-L | 88.9 |
| OpenEval - CNN/DailyMail | 25.26 | ROUGE-L | 88.2 |
| HELM Classic - bAbI | 58.39 | Exact Match (%) | 87 |
| HELM Classic - MMLU | 50.67 | Exact Match (%) | 86.4 |
Interactive version: theaggregate.ai/model?slug=llama-2-13b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.