GPT-5 Codex: benchmark results

OpenAI's Codex-specialized GPT-5 model for coding and software-agent workloads. Provider: OpenAI. Released 2025-09-23. Access: API.

Unified ELO 1671 ± 1, rank #96 of 1392 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Chess Leaderboard (Reasoning)1777Elo98.5
SnakeBench33.3TrueSkill Rating95.9
AI Chess Leaderboard (Continuation)1291Elo90.5
BenchTable69.3Total Score (%)85
WebApp1K Duo75Pass@1 (%)84.4
BenchmarkList ECI135.22Capability Index (ECI)83.4
Hack-Verifiable TextArena23.3Avg HR (self-reported)81.8
Kagi LLM Benchmark70.3Accuracy (%)81.7
Wolfram LLM Benchmarking Project47.7Correct Functionality (%)64.3
LLM Stats Score26.79LLM Stats Score (conservative rating)61
Bullshit Benchmark30.9BS Detection Rate (%)59.6
SWE-rebench44.22Resolved (%)52.1

Interactive version: theaggregate.ai/model?slug=gpt-5-codex · How It Works · Data refreshed daily, snapshot 2026-09-05.