O3 Mini — benchmark results
OpenAI's compact o3-series reasoning model (January 2025). Provider: OpenAI. Released 2025-01-31. Access: API.
Unified ELO 1661 ± 17, rank #343 of 1776 rated models, from 170 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BigCodeBench-Hard | 33.1 | Instruct pass@1 (self-reported) | 100 |
| KORGym | 0.82 | Score | 100 |
| KORGym - Control and Interaction | 0.77 | Score | 100 |
| KORGym - Spatial and Geometric | 0.94 | Score | 100 |
| LLM Stats (Multilingual MMLU) | 80.7 | Score (%) | 100 |
| PARROT - Result Consistency | 54.23 | Accuracy (%) | 100 |
| PhysicsFinals | 59.9 | Score (self-reported) | 100 |
| Smolagents - MATH | 98 | Accuracy (%) | 100 |
| U-MATH - Sequences & Series | 93.51 | Accuracy (%) | 100 |
| Vector Eval - HumanEval | 98.17 | Mean Score (%) | 100 |
| Vector Eval - MATH | 96.91 | Accuracy (%) | 100 |
| WebApp1K | 96.1 | Pass@1 (%) | 100 |
Interactive version: theaggregate.ai/model?slug=o3-mini · How the rankings work · Data refreshed daily, snapshot 2026-07-22.