GPT-5 (Medium) — benchmark results

GPT-5 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-08-07. Access: API.

Unified ELO 1838 ± 25, rank #108 of 1776 rated models, from 156 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HAL Online Mind2Web42.33Accuracy (%)100
HAL USACO69.06Accuracy (%)100
IneqMath47Overall Accuracy (self-reported)100
Step Game (Lechmazur)5.49TrueSkill μ100
SudokuBench (Single-shot, 4x4)86.74x4 Solve Rate (%)100
Translation (Lechmazur)8.69Mean Score100
AI for Education Pedagogy - Social studies90.91Accuracy (%)98.9
Aider polyglot coding leaderboard86.7Pass rate (%)98.5
Elimination Game (Lechmazur)5.97TrueSkill μ98.3
MATH Level 597.92Accuracy (%)98.1
AA MATH-50099.13Accuracy (%)98
AA Long Context Reasoning72.8Accuracy (%)96.8

Interactive version: theaggregate.ai/model?slug=gpt-5-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.