O4 Mini (Medium): benchmark results

O4 Mini evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1625 ± 1, rank #336 of 1761 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MATH-Perturb (Hard)87.1Accuracy (%)94.4
LiveCodeBench84.5Pass@1 avg (%)92.6
SEAL - VISTA51.66Score91.9
UGI - Natural Intelligence43.66NatInt Score87.8
UGI - Writing45.39Writing Score85.5
NYT Connections Older Models58.4Score (%)83.6
Generalization V1 (Lechmazur)1.8Avg Rank (lower is better)80
Roboflow Vision Evals - Visual Understanding71.64Pass Rate (%)78.9
CadEval62Score (self-reported)77.8
VPCT57.5Accuracy (%)76.9
FormationEval95Accuracy (%)74.6
LLM Chess (Saplin)240.3ELO65.2

Interactive version: theaggregate.ai/model?slug=o4-mini-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.