GPT-5 Codex — benchmark results
OpenAI's Codex-specialized GPT-5 model for coding and software-agent workloads. Provider: OpenAI. Released 2025-09-23. Access: API.
Unified ELO 1785 ± 36, rank #158 of 1776 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Chess Leaderboard (Reasoning) | 1777 | Elo | 99 |
| SnakeBench | 33.3 | TrueSkill Rating | 95.9 |
| AA-LCR | 69 | Score (self-reported) | 92.7 |
| AI Chess Leaderboard (Continuation) | 1291 | Elo | 90.9 |
| CritPt | 5.1 | Accuracy (self-reported) | 90.5 |
| BenchTable | 69.3 | Total Score (%) | 85 |
| SWE-bench Verified | 72.8 | Resolved (%) | 84.8 |
| WebApp1K Duo | 75 | Pass@1 (%) | 84.4 |
| Kagi LLM Benchmark | 70.3 | Accuracy (%) | 82.1 |
| Hack-Verifiable TextArena | 23.3 | Avg HR (self-reported) | 81.8 |
| Vals AI LiveCodeBench | 84.72 | Accuracy (%) | 79.5 |
| Wolfram LLM Benchmarking Project | 47.7 | Correct Functionality (%) | 66.7 |
Interactive version: theaggregate.ai/model?slug=gpt-5-codex · How the rankings work · Data refreshed daily, snapshot 2026-07-22.