LLaMA-7B: benchmark results
Provider: Meta. Released 2023-02-24. Access: Open.
Unified ELO 1313 ± 1, rank #1332 of 1392 rated models, from 104 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - Entity Data Imputation | 83.4 | Exact Match (%) | 84.8 |
| HELM Classic - bAbI | 53.09 | Exact Match (%) | 75.4 |
| HELM Classic - TruthfulQA | 27.98 | Exact Match (%) | 69.7 |
| MMLU-by-task - College Mathematics | 34 | Accuracy (%) | 69.2 |
| HELM Classic - CivilComments | 56.28 | Exact Match (%) | 66.7 |
| HELM Classic - IMDB | 94.7 | Exact Match (%) | 66.7 |
| HELM Classic - Entity Matching | 83.59 | Exact Match (%) | 63.6 |
| HELM Classic - NaturalQuestions Closed Book | 29.75 | F1 (%) | 60.6 |
| HELM Classic - MATH | 11.19 | Equivalent (%) | 60.3 |
| HELM Classic - MATH Chain-of-Thought | 6.1 | Equivalent (%) | 58.8 |
| HELM Classic - GSM8K | 8 | Exact Match (%) | 57.4 |
| HELM Classic - NarrativeQA | 66.91 | F1 (%) | 56.9 |
Interactive version: theaggregate.ai/model?slug=llama-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.