gpt-neox-20B — benchmark results

Provider: EleutherAI. Released 2022-02-02. Access: Open.

Unified ELO 1181 ± 9, rank #1709 of 1776 rated models, from 107 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Classic - BLiMP83.85Exact Match (%)96.8
HELM Classic - Dyck74.73Exact Match (%)92.6
HELM Classic - IMDB94.8Exact Match (%)72.7
HELM Classic - MATH14.05Equivalent (%)72.1
HELM Classic - MATH Chain-of-Thought7.06Equivalent (%)64.7
HELM Classic - Entity Matching82.03Exact Match (%)51.5
HELM Classic - Synthetic Reasoning Natural16.73F1 (%)51.5
HELM Classic - MS MARCO TREC39.82NDCG@10 (%)50
HELM Classic - NaturalQuestions Open Book59.61F1 (%)47.7
HELM Classic - GSM8K5.27Exact Match (%)47.1
HELM Classic - Synthetic Reasoning Abstract20.39Exact Match (%)42.6
HELM Classic - LegalSupport51.47Exact Match (%)41.9

Interactive version: theaggregate.ai/model?slug=gpt-neox-20b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.