O1 (2024-12-17) — benchmark results

December 17, 2024 O1 snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-12-17. Access: API.

Unified ELO 1789 ± 32, rank #151 of 1776 rated models, from 83 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AraGen84.293C3H Score (%)100
Galileo Tool Tasks - BFCL v3 Multi-Turn Missing Function94Accuracy (%)100
KOFFVQA - Korean Recognition100Score (%)100
KOFFVQA - Relationship86.67Score (%)100
LMGame-Bench 20487580Score100
SEAL - Agentic Tool Use (Enterprise)70.14Score100
SEAL - Chinese1165Score100
SEAL - Instruction Following91.96Score100
Vals AI MedQA96.52Accuracy (%)100
X-Risks Leaderboard29.093C3H Score (%)100
KOFFVQA - Recognition100Score (%)98.8
Galileo Tool Tasks - BFCL v3 Multi-Turn Base Single Function96Accuracy (%)97.2

Interactive version: theaggregate.ai/model?slug=o1-2024-12-17 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.