alpaca-7B — benchmark results

Provider: Other. Released 2023-03-18. Access: Open.

Unified ELO 1120 ± 29, rank #1753 of 1776 rated models, from 93 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - LSAT22.17Exact Match (%)79.4
HELM Classic - Entity Matching85.26Exact Match (%)78.8
HELM Classic - CivilComments56.58Exact Match (%)68.2
MMLU-by-task - TruthfulQA MC248.49Accuracy (%)67.6
HELM Classic - BoolQ77.8Exact Match (%)66.7
HELM Classic - Synthetic Reasoning Natural21.49F1 (%)64.7
HELM Classic - bAbI50.33Exact Match (%)60.9
HELM Classic - MMLU38.46Exact Match (%)57.6
HELM Classic - TruthfulQA24.31Exact Match (%)56.8
MMLU-by-task - High School Mathematics27.41Accuracy (%)54.6
HELM Classic - MATH10.43Equivalent (%)52.9
MMLU-by-task - Moral Scenarios26.15Accuracy (%)52.5

Interactive version: theaggregate.ai/model?slug=alpaca-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.