pythia-6.9B — benchmark results

Provider: EleutherAI. Released 2023-02-14. Access: Open.

Unified ELO 1141 ± 27, rank #1747 of 1776 rated models, from 45 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - Dyck65.2Exact Match (%)73.5
HELM Classic - Entity Matching82.42Exact Match (%)55.3
HELM Classic - bAbI47.9Exact Match (%)50.7
HELM Classic - LegalSupport52.15Exact Match (%)48.5
HELM Classic - MATH Chain-of-Thought4.82Equivalent (%)48.5
OpenEval - IMDb93.18Exact Match (%)47.6
ToolBench - WebShop Long0Task Score44.2
HELM Classic - MATH9.1Equivalent (%)44.1
HELM Classic - LSAT19.57Exact Match (%)43.4
ToolBench - Tabletop8.6Task Score40.7
HELM Classic - IMDB92.8Exact Match (%)40.2
ToolBench - The Cat API72Task Score36

Interactive version: theaggregate.ai/model?slug=pythia-6-9b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.