O4 Mini (2025-04-16) — benchmark results
April 16, 2025 O4 Mini snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1675 ± 10, rank #310 of 1776 rated models, from 214 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Arcadia MMMU Multiple Choice | 79.46 | Accuracy (%) | 100 |
| HELM Capabilities - Omni-MATH | 72.03 | Acc | 100 |
| HELM MedHELM - Clinical Decision Support | 78.47 | Mean win rate | 100 |
| HELM MedHELM - Mental Health | 4.98 | Jury Score | 100 |
| OpenAI HumanEval | 98.78 | Accuracy (%) | 100 |
| KOFFVQA - Document Understanding | 100 | Score (%) | 98.1 |
| EuroEval Italian NLU - Sentipolc16 | 70.34 | Sentiment classification Score (%) | 97.5 |
| KOFFVQA - Relationship | 82 | Score (%) | 97.5 |
| EuroEval Icelandic NLU - ScaLA IS | 53.35 | Linguistic acceptability Score (%) | 97.2 |
| EuroEval Italian Knowledge | 85.31 | Knowledge Average Score (%) | 97.2 |
| HELM Capabilities - GPQA | 73.54 | COT correct | 96 |
| HELM Capabilities - IFEval | 92.85 | IFEval Strict Acc | 96 |
Interactive version: theaggregate.ai/model?slug=o4-mini-2025-04-16 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.