O4 Mini (Medium) — benchmark results

O4 Mini evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1698 ± 17, rank #271 of 1776 rated models, from 32 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MATH-Perturb (Hard)87.1Accuracy (%)94.4
LiveCodeBench84.5Pass@1 avg (%)92.6
SEAL - VISTA51.66Score91.9
UGI - Natural Intelligence43.66NatInt Score88.5
UGI - Writing45.39Writing Score86.2
NYT Connections Older Models68.8Score (%)85.2
Generalization V1 (Lechmazur)1.8Avg Rank (lower is better)80
CadEval62Score (self-reported)77.8
VPCT57.5Accuracy (%)76.9
IneqMath15.5Overall Accuracy (self-reported)76.8
FormationEval95Accuracy (%)74.6
LLM Chess (Saplin)240.3ELO73.6

Interactive version: theaggregate.ai/model?slug=o4-mini-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.