O4 Mini (2025-04-16) — benchmark results

April 16, 2025 O4 Mini snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1675 ± 10, rank #310 of 1776 rated models, from 214 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Arcadia MMMU Multiple Choice79.46Accuracy (%)100
HELM Capabilities - Omni-MATH72.03Acc100
HELM MedHELM - Clinical Decision Support78.47Mean win rate100
HELM MedHELM - Mental Health4.98Jury Score100
OpenAI HumanEval98.78Accuracy (%)100
KOFFVQA - Document Understanding100Score (%)98.1
EuroEval Italian NLU - Sentipolc1670.34Sentiment classification Score (%)97.5
KOFFVQA - Relationship82Score (%)97.5
EuroEval Icelandic NLU - ScaLA IS53.35Linguistic acceptability Score (%)97.2
EuroEval Italian Knowledge85.31Knowledge Average Score (%)97.2
HELM Capabilities - GPQA73.54COT correct96
HELM Capabilities - IFEval92.85IFEval Strict Acc96

Interactive version: theaggregate.ai/model?slug=o4-mini-2025-04-16 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.