O3 Mini (High): benchmark results
O3 Mini evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-01-31. Access: API.
Unified ELO 1572 ± 1, rank #546 of 1761 rated models, from 109 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BBEH | 44.8 | Harmonic Mean (%) | 100 |
| Can LLMs Falsify? | 8.9 | Counterexample Rate (%) | 100 |
| SEAL - Agentic Tool Use (Chat) | 63.45 | Score | 100 |
| SciCode | 34.4 | Subproblem Resolve Rate (%) | 100 |
| AidanBench | 4921 | Novel Answers | 98.2 |
| ProLLM - Q&A Assistant | 98.2 | Score (%) | 94.1 |
| LiveOIBench | 60.86 | Avg Human Percentile | 93 |
| FlagEval VQA - Multi-Image Analysis | 60 | Score | 91.3 |
| ResearchCodeBench | 52.4 | Task Success Rate (%) | 87.1 |
| FlagEval VQA - Geographic Reasoning | 67.5 | Score | 87 |
| FlagEval VQA - Overall | 57.17 | Score | 87 |
| FlagEval VQA - Spatial Reasoning | 39.3 | Score | 87 |
Interactive version: theaggregate.ai/model?slug=o3-mini-high · How It Works · Data refreshed daily, snapshot 2026-09-05.