O1 (2024-12-17) — benchmark results
December 17, 2024 O1 snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-12-17. Access: API.
Unified ELO 1789 ± 32, rank #151 of 1776 rated models, from 83 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AraGen | 84.29 | 3C3H Score (%) | 100 |
| Galileo Tool Tasks - BFCL v3 Multi-Turn Missing Function | 94 | Accuracy (%) | 100 |
| KOFFVQA - Korean Recognition | 100 | Score (%) | 100 |
| KOFFVQA - Relationship | 86.67 | Score (%) | 100 |
| LMGame-Bench 2048 | 7580 | Score | 100 |
| SEAL - Agentic Tool Use (Enterprise) | 70.14 | Score | 100 |
| SEAL - Chinese | 1165 | Score | 100 |
| SEAL - Instruction Following | 91.96 | Score | 100 |
| Vals AI MedQA | 96.52 | Accuracy (%) | 100 |
| X-Risks Leaderboard | 29.09 | 3C3H Score (%) | 100 |
| KOFFVQA - Recognition | 100 | Score (%) | 98.8 |
| Galileo Tool Tasks - BFCL v3 Multi-Turn Base Single Function | 96 | Accuracy (%) | 97.2 |
Interactive version: theaggregate.ai/model?slug=o1-2024-12-17 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.