O3 Pro: benchmark results

OpenAI's pro-tier o-series reasoning model for difficult multi-step tasks. Provider: OpenAI. Released 2025-06-10. Access: API.

Unified ELO 1675 ± 1, rank #89 of 1392 rated models, from 25 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
IneqMath46Overall Accuracy (self-reported)96.3
Conceptual Reasoning Index - Consistency (ACCoRD)73.8Chance-Corrected Score (0-100)88.8
Merge-Bench46.1Equivalent text (self-reported)88.2
TrackingAI IQ Test (Mensa Norway)94.29Mensa Norway Score (%)86.5
LM Market Cap LMC Score86.7LMC Score (0-100)85.9
Kagi LLM Benchmark72.1Accuracy (%)83.8
BenchmarkList ECI133.55Capability Index (ECI)80.1
SEAL - Professional Reasoning Benchmark - Legal49.67Score77.4
AA GPQA Diamond84.55Accuracy (%)76.6
Conceptual Reasoning Index - Decision Theory (DTBench)78.22Chance-Corrected Score (0-100)76.4
SEAL - Professional Reasoning Benchmark - Finance49.08Score74.2
Artificial Analysis Intelligence Index25.75Intelligence Index73.6

Interactive version: theaggregate.ai/model?slug=o3-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.