gpt2-large — benchmark results
Provider: OpenAI. Released 2019-08-20. Access: API.
Unified ELO 921 ± 42, rank #1775 of 1776 rated models, from 100 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Energy Score (Text Generation) | 5 | Energy Score (1-5) | 79.5 |
| EuroEval Faroese NLU - ScaLA FO | 1.37 | Linguistic acceptability Score (%) | 37 |
| Open LLM Leaderboard - MuSR | 5.66 | Score | 27.5 |
| EuroEval Danish NLU - ScaLA DA | 0.58 | Linguistic acceptability Score (%) | 18.6 |
| EuroEval English NLU - SQuAD | 45.93 | Reading comprehension Score (%) | 18.1 |
| EuroEval German NLU - ScaLA DE | 2.72 | Linguistic acceptability Score (%) | 17.6 |
| EuroEval Icelandic NLU - ScaLA IS | 0 | Linguistic acceptability Score (%) | 17.4 |
| EuroEval Swedish NLU - ScaLA SV | 0.82 | Linguistic acceptability Score (%) | 16 |
| Open LLM Leaderboard - IFEval | 20.48 | Score | 12.6 |
| Open LLM Leaderboard - GPQA | 1.23 | Score | 11.8 |
| EuroEval French Knowledge | 0.41 | Knowledge Average Score (%) | 10.8 |
| EuroEval Italian NLU - ScaLA IT | 0.24 | Linguistic acceptability Score (%) | 10.4 |
Interactive version: theaggregate.ai/model?slug=gpt2-large · How the rankings work · Data refreshed daily, snapshot 2026-07-22.