O1 (High) — benchmark results

O1 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2024-12-05. Access: API.

Unified ELO 1682 ± 27, rank #304 of 1776 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PaperBench26Replication Score (%)100
PlatinumBench (MIT)0.74Avg Error Rate (%)97
ResearchCodeBench48.1Task Success Rate (%)74.2
AI for Education Visual Maths - Geometry57.69Accuracy (%)72.5
AI for Education Visual Maths61.6Accuracy (%)70.8
AI for Education Visual Maths - Number and Operations54.05Accuracy (%)68.3
AI for Education Visual Maths - Measurement81.08Accuracy (%)65.8
AI for Education Visual Maths - Algebra76.92Accuracy (%)60.8
Vals AI MATH 50090.4Accuracy (%)58
WeirdML46.08Average Score57.7
Epoch AI - ECI142.71ECI Score48.6
FutureEval8.5Unified Forecasting Score43.2

Interactive version: theaggregate.ai/model?slug=o1-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.