O1 (Medium): benchmark results
O1 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2024-12-05. Access: API.
Unified ELO 1609 ± 1, rank #391 of 1761 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI for Education Pedagogy - Social studies | 90 | Accuracy (%) | 97.7 |
| Gapminder AI Worldview | 90.9 | Correct Rate (%) | 94.1 |
| MATH Level 5 | 94.41 | Accuracy (%) | 83.3 |
| AI for Education Pedagogy - Science | 90.71 | Accuracy (%) | 82.6 |
| AI for Education Pedagogy - Secondary | 86.64 | Accuracy (%) | 80.9 |
| AI for Education Pedagogy | 86.43 | Accuracy (%) | 76.1 |
| AI for Education Pedagogy - Technology | 83.02 | Accuracy (%) | 75.8 |
| PlatinumBench (MIT) | 1.27 | Avg Error Rate (%) | 75.8 |
| AI for Education SEND | 79.36 | Accuracy (%) | 70.1 |
| LLM Chess (Saplin) | 280.1 | ELO | 67.1 |
| AI for Education Pedagogy - Primary | 87.32 | Accuracy (%) | 65.3 |
| OTIS Mock AIME 2024-25 | 73.33 | Accuracy (%) | 58.3 |
Interactive version: theaggregate.ai/model?slug=o1-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.