GPT-2 — benchmark results
Provider: OpenAI. Released 2019-02-14. Access: API.
Unified ELO 901 ± 41, rank #1776 of 1776 rated models, from 114 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| StereoSet | 72.97 | ICAT Score | 100 |
| Open LLM Leaderboard - MuSR | 18.33 | Score | 90.7 |
| AI Energy Score (Text Generation) | 5 | Energy Score (1-5) | 79.5 |
| EuroEval Faroese NLU - ScaLA FO | 0.79 | Linguistic acceptability Score (%) | 31.1 |
| EuroEval Spanish NLU - ScaLA ES | 0.01 | Linguistic acceptability Score (%) | 19.3 |
| Open LLM Leaderboard - BBH | 9.2 | Score | 19 |
| EuroEval Portuguese NLU - ScaLA PT | 0.91 | Linguistic acceptability Score (%) | 18.2 |
| EuroEval Norwegian NLU - ScaLA NN | 0.84 | Linguistic acceptability Score (%) | 16.3 |
| EuroEval Italian NLU - ScaLA IT | 1.07 | Linguistic acceptability Score (%) | 14.9 |
| EuroEval Icelandic Common Sense Reasoning | -0.05 | Common Sense Reasoning Average Score (%) | 13.3 |
| Open LLM Leaderboard - GPQA | 1.34 | Score | 13.2 |
| EuroEval Spanish Knowledge | 1.03 | Knowledge Average Score (%) | 12.7 |
Interactive version: theaggregate.ai/model?slug=gpt-2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.