O4 Mini (High) — benchmark results
O4 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1734 ± 11, rank #215 of 1776 rated models, from 130 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HLCE | 15.84 | Pass@1 (%) | 100 |
| LLM2014 Code 2025-09 - Python | 9.17 | Score | 100 |
| LiveCodeBench | 87.3 | Pass@1 avg (%) | 100 |
| TableBench | 61.69 | DP Score (%) | 100 |
| LLM2014 Logic 2025-06 | 78.97 | Median Score | 97.2 |
| AA LiveCodeBench | 85.93 | Pass@1 (%) | 96.8 |
| LLM2014 Logic 2025-05 | 80.16 | Median Score | 96.8 |
| AA MATH-500 | 98.87 | Accuracy (%) | 96.5 |
| SnakeBench | 33.7 | TrueSkill Rating | 96.5 |
| HAL TAU-bench Airline | 56 | Accuracy (%) | 96.4 |
| LLM2014 Logic 2025-04 | 80.87 | Median Score | 96.4 |
| LLM2014 Code 2025-09 - C# | 8.65 | Score | 94.7 |
Interactive version: theaggregate.ai/model?slug=o4-mini-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.