LLaMA-30B — benchmark results
Provider: Meta. Released 2023-04-04. Access: Open.
Unified ELO 1301 ± 17, rank #1533 of 1776 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Classic - LegalSupport | 63.8 | Exact Match (%) | 98.5 |
| HELM Classic - Entity Data Imputation | 84.36 | Exact Match (%) | 97.7 |
| HELM Classic - RAFT | 75.23 | Exact Match (%) | 97 |
| HELM Classic - NarrativeQA | 75.25 | F1 (%) | 96.9 |
| HELM Classic - NaturalQuestions Closed Book | 40.83 | F1 (%) | 95.5 |
| ToolBench - WebShop Short | 30.6 | Task Score | 95.3 |
| HELM Classic - bAbI | 63.36 | Exact Match (%) | 92.8 |
| MMLU-by-task - Public Relations | 70 | Accuracy (%) | 92 |
| HELM Classic - Entity Matching | 90.83 | Exact Match (%) | 90.9 |
| MMLU-by-task - Medical Genetics | 66 | Accuracy (%) | 90 |
| HELM Classic - BoolQ | 86.1 | Exact Match (%) | 89.4 |
| HELM Classic - MMLU | 53.14 | Exact Match (%) | 89.4 |
Interactive version: theaggregate.ai/model?slug=llama-30b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.