GPT-5 Codex (High) — benchmark results
GPT-5 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-09-23. Access: API.
Unified ELO 1827 ± 29, rank #118 of 1776 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| WebCoderBench - Visual Experience | 93.83 | Score (%) | 100 |
| AA AIME 2025 | 98.67 | Accuracy (%) | 99.3 |
| AA MMLU-Pro | 86.49 | Accuracy (%) | 95.6 |
| AA LiveCodeBench | 84.02 | Pass@1 (%) | 95.3 |
| WebCoderBench - Functionality Correctness | 74.82 | Score (%) | 92.3 |
| AA IFBench | 74.15 | Accuracy (%) | 92 |
| AA Omniscience - Software Engineering (SWE) - Swift | 72 | Accuracy (%) | 91.7 |
| AA Omniscience - Business | 36.7 | Accuracy (%) | 91.5 |
| AA Long Context Reasoning | 69 | Accuracy (%) | 91 |
| AA Omniscience - Law | 34.5 | Accuracy (%) | 89.3 |
| AA Omniscience - Software Engineering (SWE) - Java | 42 | Accuracy (%) | 88.8 |
| AA Omniscience - Software Engineering (SWE) - Dart | 44 | Accuracy (%) | 88.2 |
Interactive version: theaggregate.ai/model?slug=gpt-5-codex-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.