pythia-12B — benchmark results

Provider: EleutherAI. Released 2023-02-28. Access: Open.

Unified ELO 1152 ± 17, rank #1733 of 1776 rated models, from 102 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Dyck79.8Exact Match (%)97.1
HELM Classic - MATH Chain-of-Thought6.8Equivalent (%)63.2
HELM Classic - bAbI49.12Exact Match (%)56.5
MMLU-by-task - Machine Learning33.04Accuracy (%)53.1
HELM Classic - MATH10.05Equivalent (%)51.5
MMLU-by-task - High School Mathematics27.04Accuracy (%)50.5
MMLU-by-task - High School Statistics36.57Accuracy (%)45.6
HELM Classic - Entity Matching77.98Exact Match (%)45.5
MMLU-by-task - College Mathematics31Accuracy (%)45.4
ToolBench - WebShop Long0Task Score44.2
HELM Classic - IMDB93.1Exact Match (%)42.4
MMLU-by-task - Elementary Mathematics26.98Accuracy (%)39

Interactive version: theaggregate.ai/model?slug=pythia-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.