O3 Mini (Medium) — benchmark results
O3 Mini evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-01-31. Access: API.
Unified ELO 1666 ± 24, rank #330 of 1776 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ZebraLogic | 88.9 | Puzzle Accuracy (%) | 98.4 |
| AidanBench | 3368 | Novel Answers | 96.4 |
| SEAL - Agentic Tool Use (Chat) | 62.42 | Score | 93.9 |
| ProLLM - StackEval | 97.4 | Score (%) | 90.7 |
| SciCode | 33 | Subproblem Resolve Rate (%) | 88.9 |
| MATH Level 5 | 95.17 | Accuracy (%) | 86.1 |
| ProLLM - Function Calling | 91.1 | Score (%) | 86.1 |
| ProLLM - LLM-as-a-Judge | 82.4 | Score (%) | 85.7 |
| Step Game (Lechmazur) | 3.38 | TrueSkill μ | 82.4 |
| MATH-Perturb (Hard) | 82.92 | Accuracy (%) | 80.6 |
| ProLLM - Q&A Assistant | 96.4 | Score (%) | 80.4 |
| SEAL - Agentic Tool Use (Enterprise) | 64.93 | Score | 77.3 |
Interactive version: theaggregate.ai/model?slug=o3-mini-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.