O3 Mini (High) — benchmark results

O3 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-01-31. Access: API.

Unified ELO 1668 ± 18, rank #328 of 1776 rated models, from 102 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BBEH44.8Harmonic Mean (%)100
Can LLMs Falsify?8.9Counterexample Rate (%)100
SEAL - Agentic Tool Use (Chat)63.45Score100
SciCode34.4Subproblem Resolve Rate (%)100
AidanBench4921Novel Answers98.2
Defects4J48.8Defects4J Plausible @1 (self-reported)96.9
AA MATH-50098.47Accuracy (%)95
ProLLM - Q&A Assistant98.2Score (%)94.1
LiveOIBench60.86Avg Human Percentile93
ResearchCodeBench52.4Task Success Rate (%)87.1
Step Game (Lechmazur)3.64TrueSkill μ86.5
Software Engineering Arena - Model Arena1002Elo Rating86.4

Interactive version: theaggregate.ai/model?slug=o3-mini-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.