GPT-OSS-20B (Medium) — benchmark results

GPT-OSS-20B evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.

Unified ELO 1564 ± 5, rank #580 of 1776 rated models, from 296 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Polish Summarization - PSC25.35Score (%)95.9
EuroEval Slovak NLU - UNER SK65.77Named entity recognition Score (%)95.8
EuroEval Latvian NLU - Latvian Twitter Sentiment51.88Sentiment classification Score (%)95.3
EuroEval Spanish NLU - ScaLA ES41.52Linguistic acceptability Score (%)92
EuroEval Albanian NLU - MMS SQ27.91Sentiment classification Score (%)90.5
EuroEval German NLU - Sb10K58.47Sentiment classification Score (%)90.5
EuroEval Portuguese NLU - ScaLA PT39.43Linguistic acceptability Score (%)89.3
EuroEval Czech Knowledge84.67Knowledge Average Score (%)89
EuroEval Italian NLU - ScaLA IT42.67Linguistic acceptability Score (%)88.2
EuroEval English NLU - ScaLA EN58.1Linguistic acceptability Score (%)88
EuroEval French Knowledge74.85Knowledge Average Score (%)87.5
EuroEval Italian NLU61.15NLU Average Score (%)87.4

Interactive version: theaggregate.ai/model?slug=gpt-oss-20b-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.