Llama 2 7B: benchmark results
Provider: Meta. Released 2023-07-18. Access: Open.
Unified ELO 1333 ± 1, rank #1300 of 1392 rated models, from 76 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - Synthetic Reasoning Natural | 26.02 | F1 (%) | 82.4 |
| HELM Classic - bAbI | 54.97 | Exact Match (%) | 81.9 |
| HELM Classic - QuAC | 40.62 | F1 (%) | 80 |
| ToolBench - WebShop Short | 6.92 | Task Score | 79.1 |
| HELM Classic - WikiFact | 33.49 | Exact Match (%) | 75.8 |
| HELM Classic - Dyck | 65.8 | Exact Match (%) | 75 |
| HELM Classic - Entity Matching | 85.07 | Exact Match (%) | 74.2 |
| HELM Classic - MATH Chain-of-Thought | 9.91 | Equivalent (%) | 72.1 |
| ToolBench - VirtualHome | 21.49 | Task Score | 69.8 |
| HELM Classic - MMLU | 43.07 | Exact Match (%) | 68.2 |
| HELM Classic - NaturalQuestions Closed Book | 33.66 | F1 (%) | 68.2 |
| HELM Classic - NarrativeQA | 69.12 | F1 (%) | 67.7 |
Interactive version: theaggregate.ai/model?slug=llama-2-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.