O3 (2025-04-16) — benchmark results
April 16, 2025 O3 snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1722 ± 19, rank #236 of 1776 rated models, from 169 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Arena-Hard v2 | 85.9 | Win Rate (%) | 100 |
| Arena-Hard v2 (GPT-4.1 Judge) | 87 | Win Rate (%) | 100 |
| EuroEval French NLU - Allocine | 98.19 | Sentiment classification Score (%) | 100 |
| HELM Capabilities - GPQA | 75.34 | COT correct | 100 |
| HindiGen | 85.56 | 3C3H Score (%) | 100 |
| KOFFVQA - Commonsense Reasoning | 96.22 | Score (%) | 100 |
| LMGame-Bench Candy Crush | 647 | Score | 100 |
| LMGame-Bench Super Mario Bros | 3445 | Score | 100 |
| EuroEval German Knowledge | 91.87 | Knowledge Average Score (%) | 99.7 |
| EuroEval Danish NLU - Dansk | 71.31 | Named entity recognition Score (%) | 99.5 |
| EuroEval Portuguese Knowledge | 92.18 | Knowledge Average Score (%) | 99.5 |
| EuroEval French Knowledge | 87.14 | Knowledge Average Score (%) | 99.4 |
Interactive version: theaggregate.ai/model?slug=o3-2025-04-16 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.