O4 Mini: benchmark results
OpenAI's efficient o4-series reasoning model for fast math, coding, and visual tasks. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1613 ± 1, rank #224 of 1392 rated models, from 216 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MME-Reasoning | 57.2 | Overall Accuracy (%) | 100 |
| ZEROBench-Sub | 29.05 | Score (self-reported) | 97.5 |
| MMMU Benchmark | 81.6 | Validation Score | 96.8 |
| SnakeBench | 33.9 | TrueSkill Rating | 96.8 |
| TRLawBench | 84.4 | Stage 1 accuracy (self-reported) | 96.8 |
| TableBench | 60.75 | DP Score (%) | 96.8 |
| LLM Stats (AIME 2024) | 93.4 | Score (%) | 96.2 |
| VisualPuzzles | 57 | Overall Accuracy (%) | 94.6 |
| OJBench | 33.3 | Score (self-reported) | 94.1 |
| RAI-Bench - RAG Robustness (HY Factuality) | 47 | Rate (%) | 92.8 |
| LLM Stats (MathVista) | 84.3 | Score (%) | 91.9 |
| MedGUIDE | 58.9 | Weighted Accuracy (self-reported) | 91.7 |
Interactive version: theaggregate.ai/model?slug=o4-mini · How It Works · Data refreshed daily, snapshot 2026-09-05.