O3 Mini (High) — benchmark results
O3 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-01-31. Access: API.
Unified ELO 1668 ± 18, rank #328 of 1776 rated models, from 102 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BBEH | 44.8 | Harmonic Mean (%) | 100 |
| Can LLMs Falsify? | 8.9 | Counterexample Rate (%) | 100 |
| SEAL - Agentic Tool Use (Chat) | 63.45 | Score | 100 |
| SciCode | 34.4 | Subproblem Resolve Rate (%) | 100 |
| AidanBench | 4921 | Novel Answers | 98.2 |
| Defects4J | 48.8 | Defects4J Plausible @1 (self-reported) | 96.9 |
| AA MATH-500 | 98.47 | Accuracy (%) | 95 |
| ProLLM - Q&A Assistant | 98.2 | Score (%) | 94.1 |
| LiveOIBench | 60.86 | Avg Human Percentile | 93 |
| ResearchCodeBench | 52.4 | Task Success Rate (%) | 87.1 |
| Step Game (Lechmazur) | 3.64 | TrueSkill μ | 86.5 |
| Software Engineering Arena - Model Arena | 1002 | Elo Rating | 86.4 |
Interactive version: theaggregate.ai/model?slug=o3-mini-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.