O3 (2025-04-16) (Medium) — benchmark results
O3 (2025-04-16) evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1845 ± 69, rank #107 of 1776 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Fiction.LiveBench | 100 | Accuracy (%) | 100 |
| UGI - Writing | 66.83 | Writing Score | 97.3 |
| UGI - Natural Intelligence | 65.92 | NatInt Score | 96.9 |
| SEAL - EnigmaEval | 13.09 | Score | 91.9 |
| UGI Leaderboard | 47.45 | UGI Score | 84.9 |
| VPCT | 52 | Accuracy (%) | 71.8 |
| MultiNRC | 44.45 | Score (self-reported) | 65.8 |
| Wolfram LLM Benchmarking Project | 47.4 | Correct Functionality (%) | 65.4 |
| SEAL - MultiNRC | 44.45 | Score | 65.1 |
| SEAL - TutorBench | 52.76 | Score | 53.8 |
| TutorBench | 52.76 | Score (self-reported) | 50 |
| DeepResearchBench | 46.6 | Average Score | 40 |
Interactive version: theaggregate.ai/model?slug=o3-2025-04-16-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.