O1 (Medium) — benchmark results
O1 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2024-12-05. Access: API.
Unified ELO 1688 ± 34, rank #298 of 1776 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI for Education Pedagogy - Social studies | 90 | Accuracy (%) | 97.5 |
| Gapminder AI Worldview | 90.9 | Correct Rate (%) | 94.1 |
| AI for Education Pedagogy - Science | 90.71 | Accuracy (%) | 85 |
| AI for Education Pedagogy - Secondary | 86.64 | Accuracy (%) | 83.4 |
| MATH Level 5 | 94.41 | Accuracy (%) | 83.3 |
| AI for Education Pedagogy | 86.43 | Accuracy (%) | 78.6 |
| AI for Education Pedagogy - Technology | 83.02 | Accuracy (%) | 77.5 |
| PlatinumBench (MIT) | 1.27 | Avg Error Rate (%) | 75.8 |
| LLM Chess (Saplin) | 280.1 | ELO | 75 |
| AI for Education SEND | 79.36 | Accuracy (%) | 73.6 |
| AI for Education Pedagogy - Primary | 87.32 | Accuracy (%) | 69.1 |
| IneqMath | 8 | Overall Accuracy (self-reported) | 62.2 |
Interactive version: theaggregate.ai/model?slug=o1-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.