O4 Mini (High): benchmark results
O4 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.
Unified ELO 1603 ± 1, rank #425 of 1761 rated models, from 117 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HLCE | 15.84 | Pass@1 (%) | 100 |
| LLM2014 Code 2025-09 - Python | 9.17 | Score | 100 |
| LiveCodeBench | 87.3 | Pass@1 avg (%) | 100 |
| TableBench | 61.69 | DP Score (%) | 100 |
| LLM2014 Logic 2025-06 | 78.97 | Median Score | 97.2 |
| LLM2014 Logic 2025-05 | 80.16 | Median Score | 96.8 |
| SnakeBench | 33.7 | TrueSkill Rating | 96.5 |
| HAL TAU-bench Airline | 56 | Accuracy (%) | 96.4 |
| LLM2014 Logic 2025-04 | 80.87 | Median Score | 96.4 |
| FlagEval VQA - Visual Puzzles | 40.9 | Score | 95.7 |
| LLM2014 Code 2025-09 - C# | 8.65 | Score | 94.7 |
| LLM2014 Code 2025-09 - Golang | 7.35 | Score | 94.7 |
Interactive version: theaggregate.ai/model?slug=o4-mini-high · How It Works · Data refreshed daily, snapshot 2026-09-05.