GPT-2 Medium — benchmark results

Provider: OpenAI. Released 2019-05-03. Access: API.

Unified ELO 929 ± 38, rank #1774 of 1776 rated models, from 101 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Energy Score (Text Generation)5Energy Score (1-5)79.5
StereoSet71.73ICAT Score77.8
EuroEval Icelandic Common Sense Reasoning3.57Common Sense Reasoning Average Score (%)38.4
Open LLM Leaderboard - MuSR6.16Score29.6
EuroEval Spanish NLU - ScaLA ES0.73Linguistic acceptability Score (%)26.3
EuroEval Norwegian NLU - ScaLA NN1.06Linguistic acceptability Score (%)18.7
Open LLM Leaderboard - GPQA1.68Score16.5
EuroEval Portuguese NLU - ScaLA PT0.62Linguistic acceptability Score (%)16.4
Open LLM Leaderboard - IFEval22.08Score15.1
EuroEval English NLU - SQuAD38.64Reading comprehension Score (%)15
EuroEval Swedish NLU - ScaLA SV0.21Linguistic acceptability Score (%)12.8
EuroEval German Knowledge0.79Knowledge Average Score (%)11.7

Interactive version: theaggregate.ai/model?slug=gpt-2-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.