GPT-OSS-20B (Medium) — benchmark results
GPT-OSS-20B evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1564 ± 5, rank #580 of 1776 rated models, from 296 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Polish Summarization - PSC | 25.35 | Score (%) | 95.9 |
| EuroEval Slovak NLU - UNER SK | 65.77 | Named entity recognition Score (%) | 95.8 |
| EuroEval Latvian NLU - Latvian Twitter Sentiment | 51.88 | Sentiment classification Score (%) | 95.3 |
| EuroEval Spanish NLU - ScaLA ES | 41.52 | Linguistic acceptability Score (%) | 92 |
| EuroEval Albanian NLU - MMS SQ | 27.91 | Sentiment classification Score (%) | 90.5 |
| EuroEval German NLU - Sb10K | 58.47 | Sentiment classification Score (%) | 90.5 |
| EuroEval Portuguese NLU - ScaLA PT | 39.43 | Linguistic acceptability Score (%) | 89.3 |
| EuroEval Czech Knowledge | 84.67 | Knowledge Average Score (%) | 89 |
| EuroEval Italian NLU - ScaLA IT | 42.67 | Linguistic acceptability Score (%) | 88.2 |
| EuroEval English NLU - ScaLA EN | 58.1 | Linguistic acceptability Score (%) | 88 |
| EuroEval French Knowledge | 74.85 | Knowledge Average Score (%) | 87.5 |
| EuroEval Italian NLU | 61.15 | NLU Average Score (%) | 87.4 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-20b-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.