O4 Mini (High) — benchmark results

O4 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1734 ± 11, rank #215 of 1776 rated models, from 130 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HLCE15.84Pass@1 (%)100
LLM2014 Code 2025-09 - Python9.17Score100
LiveCodeBench87.3Pass@1 avg (%)100
TableBench61.69DP Score (%)100
LLM2014 Logic 2025-0678.97Median Score97.2
AA LiveCodeBench85.93Pass@1 (%)96.8
LLM2014 Logic 2025-0580.16Median Score96.8
AA MATH-50098.87Accuracy (%)96.5
SnakeBench33.7TrueSkill Rating96.5
HAL TAU-bench Airline56Accuracy (%)96.4
LLM2014 Logic 2025-0480.87Median Score96.4
LLM2014 Code 2025-09 - C#8.65Score94.7

Interactive version: theaggregate.ai/model?slug=o4-mini-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.