O3 Mini: benchmark results
OpenAI's compact o3-series reasoning model (January 2025). Provider: OpenAI. Released 2025-01-31. Access: API.
Unified ELO 1602 ± 1, rank #260 of 1392 rated models, from 189 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| KORGym | 0.82 | Score | 100 |
| KORGym - Control and Interaction | 0.77 | Score | 100 |
| KORGym - Spatial and Geometric | 0.94 | Score | 100 |
| LLM Stats (Multilingual MMLU) | 80.7 | Score (%) | 100 |
| PARROT - Result Consistency | 54.23 | Accuracy (%) | 100 |
| Smolagents - MATH | 98 | Accuracy (%) | 100 |
| U-MATH - Sequences & Series | 93.51 | Accuracy (%) | 100 |
| Vector Eval - HumanEval | 98.17 | Mean Score (%) | 100 |
| WebApp1K | 96.1 | Pass@1 (%) | 100 |
| U-MATH - Precalculus | 94.38 | Accuracy (%) | 97 |
| PhysicsFinals | 59.9 | Score (self-reported) | 96.9 |
| SEAL - Coding | 1137 | Score | 96.3 |
Interactive version: theaggregate.ai/model?slug=o3-mini · How It Works · Data refreshed daily, snapshot 2026-09-05.