O1 (High): benchmark results

O1 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2024-12-05. Access: API.

Unified ELO 1601 ± 1, rank #433 of 1761 rated models, from 15 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PaperBench26Replication Score (%)100
PlatinumBench (MIT)0.74Avg Error Rate (%)97
ResearchCodeBench48.1Task Success Rate (%)74.2
FutureEval8.23Unified Forecasting Score66.3
AI for Education Visual Maths - Geometry57.69Accuracy (%)62.8
AI for Education Visual Maths61.6Accuracy (%)61.5
AI for Education Visual Maths - Number and Operations54.05Accuracy (%)59.5
ComplexFuncBench47.6Score (self-reported)57.9
AI for Education Visual Maths - Measurement81.08Accuracy (%)54.7
AI for Education Visual Maths - Algebra76.92Accuracy (%)54.1
WeirdML46.08Average Score52.2
FrontierMath - Tiers 1-39.31Accuracy (%, 290 problems)42.9

Interactive version: theaggregate.ai/model?slug=o1-high · How It Works · Data refreshed daily, snapshot 2026-09-05.