O3 (Medium): benchmark results
O3 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1639 ± 1, rank #284 of 1761 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Gapminder AI Worldview | 93.9 | Correct Rate (%) | 100 |
| HAL AssistantBench | 38.81 | Accuracy (%) | 100 |
| HAL ScienceAgentBench | 33.33 | Accuracy (%) | 100 |
| Step Game (Lechmazur) | 5.32 | TrueSkill μ | 98.6 |
| HAL SciCode | 9.23 | Accuracy (%) | 96.7 |
| NYT Connections Older Models | 63 | Score (%) | 87.3 |
| LisanBench | 0.21 | Mean Path Length / Current Maximum | 86.3 |
| HAL TAU-bench Airline | 54 | Accuracy (%) | 82.1 |
| LLM Chess (Saplin) | 777.6 | ELO | 82 |
| EnigmaEval | 13.09 | Score (self-reported) | 80 |
| SEAL - VISTA | 49.59 | Score | 79 |
| SEAL - Humanity's Last Exam (Text Only) | 19.78 | Score | 76.7 |
Interactive version: theaggregate.ai/model?slug=o3-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.