O3 Mini — benchmark results

OpenAI's compact o3-series reasoning model (January 2025). Provider: OpenAI. Released 2025-01-31. Access: API.

Unified ELO 1661 ± 17, rank #343 of 1776 rated models, from 170 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BigCodeBench-Hard33.1Instruct pass@1 (self-reported)100
KORGym0.82Score100
KORGym - Control and Interaction0.77Score100
KORGym - Spatial and Geometric0.94Score100
LLM Stats (Multilingual MMLU)80.7Score (%)100
PARROT - Result Consistency54.23Accuracy (%)100
PhysicsFinals59.9Score (self-reported)100
Smolagents - MATH98Accuracy (%)100
U-MATH - Sequences & Series93.51Accuracy (%)100
Vector Eval - HumanEval98.17Mean Score (%)100
Vector Eval - MATH96.91Accuracy (%)100
WebApp1K96.1Pass@1 (%)100

Interactive version: theaggregate.ai/model?slug=o3-mini · How the rankings work · Data refreshed daily, snapshot 2026-07-22.