LLaMA-7B — benchmark results

Provider: Meta. Released 2023-02-24. Access: Open.

Unified ELO 1205 ± 14, rank #1686 of 1776 rated models, from 110 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Entity Data Imputation83.4Exact Match (%)84.8
HELM Classic - bAbI53.09Exact Match (%)75.4
OpenEval - CNN/DailyMail23.52ROUGE-L70.6
HELM Classic - TruthfulQA27.98Exact Match (%)69.7
MMLU-by-task - College Mathematics34Accuracy (%)69.2
HELM Classic - CivilComments56.28Exact Match (%)66.7
HELM Classic - IMDB94.7Exact Match (%)66.7
HELM Classic - Entity Matching83.59Exact Match (%)63.6
HELM Classic - NaturalQuestions Closed Book29.75F1 (%)60.6
HELM Classic - MATH11.19Equivalent (%)60.3
HELM Classic - MATH Chain-of-Thought6.1Equivalent (%)58.8
OpenEval - XSum23.11ROUGE-L58.3

Interactive version: theaggregate.ai/model?slug=llama-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.