O3 Mini (2025-01-31) (High) — benchmark results
O3 Mini (2025-01-31) evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-01-31. Access: API.
Unified ELO 1811 ± 66, rank #130 of 1776 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ZebraLogic | 91.7 | Puzzle Accuracy (%) | 100 |
| RepairBench | 46.4 | Plausible@1 (%) | 97.4 |
| MATH Level 5 | 96.49 | Accuracy (%) | 90.7 |
| Arena-Hard v2 (GPT-4.1 Judge) | 64.8 | Win Rate (%) | 88.9 |
| Arena-Hard v2 | 66.1 | Win Rate (%) | 81.5 |
| SuperGPQA | 55.22 | Accuracy (%) | 73.5 |
| OTIS Mock AIME 2024-25 | 76.94 | Accuracy (%) | 59.6 |
| LiveCodeBench | 77.7 | Pass@1 avg (%) | 55.6 |
| Arena-Hard Creative Writing | 43 | Win Rate (%) | 40.7 |
| Chess Puzzles (Epoch AI) | 17 | Accuracy (%) | 24.1 |
Interactive version: theaggregate.ai/model?slug=o3-mini-2025-01-31-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.