O1 Preview (2024-09-12) — benchmark results
September 12, 2024 O1 Preview snapshot, tracked when sources report the dated API model. Provider: OpenAI. Released 2024-09-12. Access: API.
Unified ELO 1666 ± 50, rank #331 of 1776 rated models, from 9 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ZebraLogic | 71.4 | Puzzle Accuracy (%) | 91.8 |
| Wolfram LLM Benchmarking Project | 52.2 | Correct Functionality (%) | 75.8 |
| Vals AI MedQA | 93.01 | Accuracy (%) | 73.4 |
| MATH Level 5 | 81.65 | Accuracy (%) | 65.7 |
| CRUST-bench | 15 | Pass@1 Test Success (%) | 64.3 |
| BIRD-CRITIC | 33.33 | Score | 60 |
| OTIS Mock AIME 2024-25 | 31.11 | Accuracy (%) | 34.9 |
| WritingBench | 63.44 | Score (self-reported) | 30.8 |
| SpeechMap Compliance | 35.8 | % Requests Completed | 12.4 |
Interactive version: theaggregate.ai/model?slug=o1-preview-2024-09-12 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.