O1 (2024-12-17): benchmark results

December 17, 2024 O1 snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-12-17. Access: API.

Unified ELO 1640 ± 1, rank #148 of 1392 rated models, from 123 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AraGen84.293C3H Score (%)100
Galileo Tool Tasks - BFCL v3 Multi-Turn Missing Function94Accuracy (%)100
KOFFVQA - Korean Recognition100Score (%)100
KOFFVQA - Relationship86.67Score (%)100
LMGame-Bench 20487580Score100
SEAL - Agentic Tool Use (Enterprise)70.14Score100
SEAL - Chinese1165Score100
SEAL - Instruction Following91.96Score100
X-Risks Leaderboard29.093C3H Score (%)100
KOFFVQA - Recognition100Score (%)98.8
Galileo Tool Tasks - BFCL v3 Multi-Turn Base Single Function96Accuracy (%)97.2
Galileo Tool Tasks - xLAM Single Tool Single Call100Accuracy (%)97.2

Interactive version: theaggregate.ai/model?slug=o1-2024-12-17 · How It Works · Data refreshed daily, snapshot 2026-09-05.