O1 — benchmark results
OpenAI's first-generation o-series reasoning model for difficult multi-step tasks. Provider: OpenAI. Released 2024-12-05. Access: API.
Unified ELO 1696 ± 14, rank #277 of 1776 rated models, from 223 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AidanBench | 6054 | Novel Answers | 100 |
| BlueBench - Knowledge | 71.43 | Score (%) | 100 |
| BlueBench - Reasoning | 79 | Score (%) | 100 |
| CRMArena - BRI | 74.8 | BRI Score (%) | 100 |
| CRMArena - HTU | 68.5 | HTU Score (%) | 100 |
| CRMArena - NED | 60 | NED Score (%) | 100 |
| CRMArena - Overall | 64.3 | Overall Score (%) | 100 |
| CRMArena - TII | 99.2 | TII Score (%) | 100 |
| ClinPivot | 69.32 | ClinPivot (self-reported) | 100 |
| DROP | 90.2 | F1 Score | 100 |
| KernelBench | 1 | Avg Speedup vs PyTorch | 100 |
| LLM2014 Logic 2024-12 | 94.58 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=o1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.