alpaca-7B: benchmark results

Provider: Other. Released 2023-03-18. Access: Open.

Unified ELO 1268 ± 1, rank #1376 of 1392 rated models, from 89 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - LSAT22.17Exact Match (%)79.4
HELM Classic - Entity Matching85.26Exact Match (%)78.8
HELM Classic - CivilComments56.58Exact Match (%)68.2
MMLU-by-task - TruthfulQA MC248.49Accuracy (%)67.6
HELM Classic - BoolQ77.8Exact Match (%)66.7
HELM Classic - Synthetic Reasoning Natural21.49F1 (%)64.7
HELM Classic - bAbI50.33Exact Match (%)60.9
HELM Classic - MMLU38.46Exact Match (%)57.6
InstructEval - Problem Solving32.4Average (%, MMLU/BBH/DROP/CRASS/HumanEval)57.1
HELM Classic - TruthfulQA24.31Exact Match (%)56.8
MMLU-by-task - High School Mathematics27.41Accuracy (%)54.6
HELM Classic - MATH10.43Equivalent (%)52.9

Interactive version: theaggregate.ai/model?slug=alpaca-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.