O3 (Medium): benchmark results

O3 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1639 ± 1, rank #284 of 1761 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Gapminder AI Worldview93.9Correct Rate (%)100
HAL AssistantBench38.81Accuracy (%)100
HAL ScienceAgentBench33.33Accuracy (%)100
Step Game (Lechmazur)5.32TrueSkill μ98.6
HAL SciCode9.23Accuracy (%)96.7
NYT Connections Older Models63Score (%)87.3
LisanBench0.21Mean Path Length / Current Maximum86.3
HAL TAU-bench Airline54Accuracy (%)82.1
LLM Chess (Saplin)777.6ELO82
EnigmaEval13.09Score (self-reported)80
SEAL - VISTA49.59Score79
SEAL - Humanity's Last Exam (Text Only)19.78Score76.7

Interactive version: theaggregate.ai/model?slug=o3-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.