O3 Pro: benchmark results
OpenAI's pro-tier o-series reasoning model for difficult multi-step tasks. Provider: OpenAI. Released 2025-06-10. Access: API.
Unified ELO 1675 ± 1, rank #89 of 1392 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| IneqMath | 46 | Overall Accuracy (self-reported) | 96.3 |
| Conceptual Reasoning Index - Consistency (ACCoRD) | 73.8 | Chance-Corrected Score (0-100) | 88.8 |
| Merge-Bench | 46.1 | Equivalent text (self-reported) | 88.2 |
| TrackingAI IQ Test (Mensa Norway) | 94.29 | Mensa Norway Score (%) | 86.5 |
| LM Market Cap LMC Score | 86.7 | LMC Score (0-100) | 85.9 |
| Kagi LLM Benchmark | 72.1 | Accuracy (%) | 83.8 |
| BenchmarkList ECI | 133.55 | Capability Index (ECI) | 80.1 |
| SEAL - Professional Reasoning Benchmark - Legal | 49.67 | Score | 77.4 |
| AA GPQA Diamond | 84.55 | Accuracy (%) | 76.6 |
| Conceptual Reasoning Index - Decision Theory (DTBench) | 78.22 | Chance-Corrected Score (0-100) | 76.4 |
| SEAL - Professional Reasoning Benchmark - Finance | 49.08 | Score | 74.2 |
| Artificial Analysis Intelligence Index | 25.75 | Intelligence Index | 73.6 |
Interactive version: theaggregate.ai/model?slug=o3-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.