O4 Mini (Medium): benchmark results
O4 Mini evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1625 ± 1, rank #336 of 1761 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MATH-Perturb (Hard) | 87.1 | Accuracy (%) | 94.4 |
| LiveCodeBench | 84.5 | Pass@1 avg (%) | 92.6 |
| SEAL - VISTA | 51.66 | Score | 91.9 |
| UGI - Natural Intelligence | 43.66 | NatInt Score | 87.8 |
| UGI - Writing | 45.39 | Writing Score | 85.5 |
| NYT Connections Older Models | 58.4 | Score (%) | 83.6 |
| Generalization V1 (Lechmazur) | 1.8 | Avg Rank (lower is better) | 80 |
| Roboflow Vision Evals - Visual Understanding | 71.64 | Pass Rate (%) | 78.9 |
| CadEval | 62 | Score (self-reported) | 77.8 |
| VPCT | 57.5 | Accuracy (%) | 76.9 |
| FormationEval | 95 | Accuracy (%) | 74.6 |
| LLM Chess (Saplin) | 240.3 | ELO | 65.2 |
Interactive version: theaggregate.ai/model?slug=o4-mini-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.