O3 Mini (Medium) — benchmark results

O3 Mini evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-01-31. Access: API.

Unified ELO 1666 ± 24, rank #330 of 1776 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ZebraLogic88.9Puzzle Accuracy (%)98.4
AidanBench3368Novel Answers96.4
SEAL - Agentic Tool Use (Chat)62.42Score93.9
ProLLM - StackEval97.4Score (%)90.7
SciCode33Subproblem Resolve Rate (%)88.9
MATH Level 595.17Accuracy (%)86.1
ProLLM - Function Calling91.1Score (%)86.1
ProLLM - LLM-as-a-Judge82.4Score (%)85.7
Step Game (Lechmazur)3.38TrueSkill μ82.4
MATH-Perturb (Hard)82.92Accuracy (%)80.6
ProLLM - Q&A Assistant96.4Score (%)80.4
SEAL - Agentic Tool Use (Enterprise)64.93Score77.3

Interactive version: theaggregate.ai/model?slug=o3-mini-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.