LLaMA-7B: benchmark results

Provider: Meta. Released 2023-02-24. Access: Open.

Unified ELO 1313 ± 1, rank #1332 of 1392 rated models, from 104 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Entity Data Imputation83.4Exact Match (%)84.8
HELM Classic - bAbI53.09Exact Match (%)75.4
HELM Classic - TruthfulQA27.98Exact Match (%)69.7
MMLU-by-task - College Mathematics34Accuracy (%)69.2
HELM Classic - CivilComments56.28Exact Match (%)66.7
HELM Classic - IMDB94.7Exact Match (%)66.7
HELM Classic - Entity Matching83.59Exact Match (%)63.6
HELM Classic - NaturalQuestions Closed Book29.75F1 (%)60.6
HELM Classic - MATH11.19Equivalent (%)60.3
HELM Classic - MATH Chain-of-Thought6.1Equivalent (%)58.8
HELM Classic - GSM8K8Exact Match (%)57.4
HELM Classic - NarrativeQA66.91F1 (%)56.9

Interactive version: theaggregate.ai/model?slug=llama-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.