O1 (High): benchmark results
O1 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2024-12-05. Access: API.
Unified ELO 1601 ± 1, rank #433 of 1761 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PaperBench | 26 | Replication Score (%) | 100 |
| PlatinumBench (MIT) | 0.74 | Avg Error Rate (%) | 97 |
| ResearchCodeBench | 48.1 | Task Success Rate (%) | 74.2 |
| FutureEval | 8.23 | Unified Forecasting Score | 66.3 |
| AI for Education Visual Maths - Geometry | 57.69 | Accuracy (%) | 62.8 |
| AI for Education Visual Maths | 61.6 | Accuracy (%) | 61.5 |
| AI for Education Visual Maths - Number and Operations | 54.05 | Accuracy (%) | 59.5 |
| ComplexFuncBench | 47.6 | Score (self-reported) | 57.9 |
| AI for Education Visual Maths - Measurement | 81.08 | Accuracy (%) | 54.7 |
| AI for Education Visual Maths - Algebra | 76.92 | Accuracy (%) | 54.1 |
| WeirdML | 46.08 | Average Score | 52.2 |
| FrontierMath - Tiers 1-3 | 9.31 | Accuracy (%, 290 problems) | 42.9 |
Interactive version: theaggregate.ai/model?slug=o1-high · How It Works · Data refreshed daily, snapshot 2026-09-05.