GPT-OSS-20B (High) — benchmark results
GPT-OSS-20B evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1521 ± 10, rank #718 of 1776 rated models, from 394 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Latvian NLU - Latvian Twitter Sentiment | 51.26 | Sentiment classification Score (%) | 93.5 |
| AA LiveCodeBench | 77.67 | Pass@1 (%) | 89.3 |
| EuroEval Spanish NLU - ScaLA ES | 39.3 | Linguistic acceptability Score (%) | 89.1 |
| AA AIME 2025 | 89.33 | Accuracy (%) | 88.5 |
| EuroEval Portuguese NLU - ScaLA PT | 35.96 | Linguistic acceptability Score (%) | 88 |
| EuroEval English NLU - ScaLA EN | 58.02 | Linguistic acceptability Score (%) | 87.7 |
| Medmarks - MedHallu Medium | 65.83 | Score (%) | 87.1 |
| EuroEval Italian NLU - ScaLA IT | 38.45 | Linguistic acceptability Score (%) | 85 |
| EuroEval German NLU - Sb10K | 57.44 | Sentiment classification Score (%) | 84.6 |
| Medmarks - SCTpublic | 68.35 | Score (%) | 84.3 |
| EuroEval Slovak NLU - UNER SK | 60.29 | Named entity recognition Score (%) | 83.3 |
| EuroEval Slovene Knowledge | 64.23 | Knowledge Average Score (%) | 83.3 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-20b-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.