O3 Mini: benchmark results

OpenAI's compact o3-series reasoning model (January 2025). Provider: OpenAI. Released 2025-01-31. Access: API.

Unified ELO 1602 ± 1, rank #260 of 1392 rated models, from 189 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
KORGym0.82Score100
KORGym - Control and Interaction0.77Score100
KORGym - Spatial and Geometric0.94Score100
LLM Stats (Multilingual MMLU)80.7Score (%)100
PARROT - Result Consistency54.23Accuracy (%)100
Smolagents - MATH98Accuracy (%)100
U-MATH - Sequences & Series93.51Accuracy (%)100
Vector Eval - HumanEval98.17Mean Score (%)100
WebApp1K96.1Pass@1 (%)100
U-MATH - Precalculus94.38Accuracy (%)97
PhysicsFinals59.9Score (self-reported)96.9
SEAL - Coding1137Score96.3

Interactive version: theaggregate.ai/model?slug=o3-mini · How It Works · Data refreshed daily, snapshot 2026-09-05.