GPT-5 Codex: benchmark results
OpenAI's Codex-specialized GPT-5 model for coding and software-agent workloads. Provider: OpenAI. Released 2025-09-23. Access: API.
Unified ELO 1671 ± 1, rank #96 of 1392 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Chess Leaderboard (Reasoning) | 1777 | Elo | 98.5 |
| SnakeBench | 33.3 | TrueSkill Rating | 95.9 |
| AI Chess Leaderboard (Continuation) | 1291 | Elo | 90.5 |
| BenchTable | 69.3 | Total Score (%) | 85 |
| WebApp1K Duo | 75 | Pass@1 (%) | 84.4 |
| BenchmarkList ECI | 135.22 | Capability Index (ECI) | 83.4 |
| Hack-Verifiable TextArena | 23.3 | Avg HR (self-reported) | 81.8 |
| Kagi LLM Benchmark | 70.3 | Accuracy (%) | 81.7 |
| Wolfram LLM Benchmarking Project | 47.7 | Correct Functionality (%) | 64.3 |
| LLM Stats Score | 26.79 | LLM Stats Score (conservative rating) | 61 |
| Bullshit Benchmark | 30.9 | BS Detection Rate (%) | 59.6 |
| SWE-rebench | 44.22 | Resolved (%) | 52.1 |
Interactive version: theaggregate.ai/model?slug=gpt-5-codex · How It Works · Data refreshed daily, snapshot 2026-09-05.