GPT-2 Large — benchmark results

Provider: OpenAI. Released 2019-08-20. Access: Open.

Unified ELO 1019 ± 27, rank #1803 of 1806 rated models, from 101 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Energy Score (Text Generation)5Energy Score (1-5)79.5
StereoSet70.54ICAT Score60
EuroEval Faroese NLU - ScaLA FO1.37Linguistic acceptability Score (%)37
Open LLM Leaderboard - MuSR5.66Score27.5
EuroEval Danish NLU - ScaLA DA0.58Linguistic acceptability Score (%)18.6
EuroEval English NLU - SQuAD45.93Reading comprehension Score (%)18.1
EuroEval German NLU - ScaLA DE2.72Linguistic acceptability Score (%)17.6
EuroEval Icelandic NLU - ScaLA IS0Linguistic acceptability Score (%)17.4
EuroEval Swedish NLU - ScaLA SV0.82Linguistic acceptability Score (%)16
Open LLM Leaderboard - IFEval20.48Score12.6
Open LLM Leaderboard - GPQA1.23Score11.8
EuroEval French Knowledge0.41Knowledge Average Score (%)10.8

Interactive version: theaggregate.ai/model?slug=gpt-2-large · How It Works · Data refreshed daily, snapshot 2026-08-07.