LLaMA-65B: benchmark results
Provider: Meta. Released 2023-02-24. Access: Open.
Unified ELO 1414 ± 1, rank #1085 of 1392 rated models, from 124 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ToolBench - WebShop Short | 41.2 | Task Score | 100 |
| MMLU-by-task - Computer Security | 79 | Accuracy (%) | 99.2 |
| MMLU-by-task - Public Relations | 74.55 | Accuracy (%) | 99.1 |
| HELM | 90.83 | Mean win rate (self-reported) | 98.5 |
| HELM Classic - IMDB | 96.2 | Exact Match (%) | 98.5 |
| HELM Classic - NarrativeQA | 75.49 | F1 (%) | 98.5 |
| HELM Classic - NaturalQuestions Closed Book | 43.14 | F1 (%) | 98.5 |
| HELM Classic - WikiFact | 42.08 | Exact Match (%) | 98.5 |
| MMLU-by-task - High School Statistics | 60.65 | Accuracy (%) | 98.1 |
| HELM Classic - MMLU | 58.37 | Exact Match (%) | 97 |
| MMLU-by-task - Virology | 54.22 | Accuracy (%) | 96.7 |
| MMLU-by-task - Management | 82.52 | Accuracy (%) | 96.4 |
Interactive version: theaggregate.ai/model?slug=llama-65b · How It Works · Data refreshed daily, snapshot 2026-09-05.