O4 Mini (High): benchmark results

O4 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-04-16. Access: API.

Unified ELO 1603 ± 1, rank #425 of 1761 rated models, from 117 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HLCE15.84Pass@1 (%)100
LLM2014 Code 2025-09 - Python9.17Score100
LiveCodeBench87.3Pass@1 avg (%)100
TableBench61.69DP Score (%)100
LLM2014 Logic 2025-0678.97Median Score97.2
LLM2014 Logic 2025-0580.16Median Score96.8
SnakeBench33.7TrueSkill Rating96.5
HAL TAU-bench Airline56Accuracy (%)96.4
LLM2014 Logic 2025-0480.87Median Score96.4
FlagEval VQA - Visual Puzzles40.9Score95.7
LLM2014 Code 2025-09 - C#8.65Score94.7
LLM2014 Code 2025-09 - Golang7.35Score94.7

Interactive version: theaggregate.ai/model?slug=o4-mini-high · How It Works · Data refreshed daily, snapshot 2026-09-05.