pythia-6.9B: benchmark results

Provider: EleutherAI. Released 2023-02-14. Access: Open.

Unified ELO 1256 ± 1, rank #1384 of 1392 rated models, from 45 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Dyck65.2Exact Match (%)73.5
HELM Classic - Entity Matching82.42Exact Match (%)55.3
HELM Classic - bAbI47.9Exact Match (%)50.7
OpenEval - IMDb93.18Exact Match (%)50
HELM Classic - LegalSupport52.15Exact Match (%)48.5
HELM Classic - MATH Chain-of-Thought4.82Equivalent (%)48.5
ToolBench - WebShop Long0Task Score44.2
HELM Classic - MATH9.1Equivalent (%)44.1
HELM Classic - LSAT19.57Exact Match (%)43.4
ToolBench - Tabletop8.6Task Score40.7
HELM Classic - IMDB92.8Exact Match (%)40.2
ToolBench - The Cat API72Task Score36

Interactive version: theaggregate.ai/model?slug=pythia-6-9b · How It Works · Data refreshed daily, snapshot 2026-09-05.