O3 (2025-04-16) (High): benchmark results
O3 (2025-04-16) evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1655 ± 1, rank #228 of 1761 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Writing | 66.75 | Writing Score | 96.5 |
| UGI - Natural Intelligence | 66.19 | NatInt Score | 96.4 |
| MATH Level 5 | 97.77 | Accuracy (%) | 95.4 |
| Vals AI MedQA | 96.06 | Accuracy (%) | 94.1 |
| Vals AI MATH 500 | 94.6 | Accuracy (%) | 89.4 |
| UGI Leaderboard | 46.52 | UGI Score | 83.5 |
| Chess Puzzles (Epoch AI) | 34 | Accuracy (%) | 83.3 |
| Epoch AI - Mystery Game Puzzles | 29 | Score | 79.4 |
| Vals AI TaxEval v2 | 74.57 | Accuracy (%) | 77.1 |
| Conceptual Reasoning Index - Argument Evaluation (LMCA) | 39.7 | Chance-Corrected Score (0-100) | 77 |
| SimpleQA Verified | 49.4 | Accuracy (%) | 74 |
| MedCode | 47.29 | Score (self-reported) | 73 |
Interactive version: theaggregate.ai/model?slug=o3-2025-04-16-high · How It Works · Data refreshed daily, snapshot 2026-09-05.