LLaMA-65B — benchmark results
Provider: Meta. Released 2023-02-24. Access: Open.
Unified ELO 1353 ± 11, rank #1400 of 1776 rated models, from 129 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ToolBench - WebShop Short | 41.2 | Task Score | 100 |
| MMLU-by-task - Computer Security | 79 | Accuracy (%) | 99.2 |
| MMLU-by-task - Public Relations | 74.55 | Accuracy (%) | 99.1 |
| HELM Classic - IMDB | 96.2 | Exact Match (%) | 98.5 |
| HELM Classic - NarrativeQA | 75.49 | F1 (%) | 98.5 |
| HELM Classic - NaturalQuestions Closed Book | 43.14 | F1 (%) | 98.5 |
| HELM Classic - WikiFact | 42.08 | Exact Match (%) | 98.5 |
| HELM | 90.83 | Mean win rate (self-reported) | 98.4 |
| MMLU-by-task - High School Statistics | 60.65 | Accuracy (%) | 98.1 |
| HELM Classic - MMLU | 58.37 | Exact Match (%) | 97 |
| MMLU-by-task - Virology | 54.22 | Accuracy (%) | 96.7 |
| MMLU-by-task - Management | 82.52 | Accuracy (%) | 96.4 |
Interactive version: theaggregate.ai/model?slug=llama-65b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.