O3 Pro — benchmark results
OpenAI's pro-tier o-series reasoning model for difficult multi-step tasks. Provider: OpenAI. Released 2025-06-10. Access: API.
Unified ELO 1757 ± 25, rank #183 of 1776 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| TrackingAI IQ Test (Mensa Norway) | 94.29 | Mensa Norway Score (%) | 93.9 |
| Merge-Bench | 46.1 | Equivalent text (self-reported) | 88.2 |
| SEAL - Professional Reasoning Benchmark - Legal | 49.67 | Score | 85.7 |
| TrackingAI IQ Test | 90.2 | IQ Test Score (%) | 85 |
| Kagi LLM Benchmark | 72.1 | Accuracy (%) | 84.3 |
| AA GPQA Diamond | 84.55 | Accuracy (%) | 82.6 |
| SEAL - Professional Reasoning Benchmark - Finance | 49.08 | Score | 82.1 |
| TutorBench | 54.62 | Score (self-reported) | 81.8 |
| Artificial Analysis Intelligence Index | 32.52 | Intelligence Index | 79.2 |
| Ducky Bench (Saxo Frog) | 1158 | ELO | 70.8 |
| TrackingAI IQ Test (Offline) | 68.75 | Offline IQ Score (%) | 66.2 |
| NarrativeWorldBench | 79 | Plot-Beat F1 (h=50) (self-reported) | 61.1 |
Interactive version: theaggregate.ai/model?slug=o3-pro · How the rankings work · Data refreshed daily, snapshot 2026-07-22.