O4 Mini (Medium) — benchmark results
O4 Mini evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1698 ± 17, rank #271 of 1776 rated models, from 32 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MATH-Perturb (Hard) | 87.1 | Accuracy (%) | 94.4 |
| LiveCodeBench | 84.5 | Pass@1 avg (%) | 92.6 |
| SEAL - VISTA | 51.66 | Score | 91.9 |
| UGI - Natural Intelligence | 43.66 | NatInt Score | 88.5 |
| UGI - Writing | 45.39 | Writing Score | 86.2 |
| NYT Connections Older Models | 68.8 | Score (%) | 85.2 |
| Generalization V1 (Lechmazur) | 1.8 | Avg Rank (lower is better) | 80 |
| CadEval | 62 | Score (self-reported) | 77.8 |
| VPCT | 57.5 | Accuracy (%) | 76.9 |
| IneqMath | 15.5 | Overall Accuracy (self-reported) | 76.8 |
| FormationEval | 95 | Accuracy (%) | 74.6 |
| LLM Chess (Saplin) | 240.3 | ELO | 73.6 |
Interactive version: theaggregate.ai/model?slug=o4-mini-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.