GPT-5.2 Codex — benchmark results
OpenAI's coding-specialized GPT-5.2 Codex model for software tasks. Provider: OpenAI. Released 2026-01-14. Access: API.
Unified ELO 1760 ± 20, rank #180 of 1776 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA-LCR | 75.7 | Score (self-reported) | 100 |
| Vals AI LiveCodeBench | 87.99 | Accuracy (%) | 98.4 |
| CritPt | 8.7 | Accuracy (self-reported) | 93.3 |
| SnakeBench | 32.4 | TrueSkill Rating | 91.6 |
| AI Chess Leaderboard (Reasoning) | 1303 | Elo | 90 |
| AISI Cyber CTF | 72 | Success Rate (%) | 89.3 |
| AI Chess Leaderboard (Continuation) | 1151 | Elo | 87.2 |
| Terminal-Bench 2.0 | 66.5 | Accuracy (%) | 86.2 |
| Epoch AI - Apex Agents | 27.6 | Score | 72.9 |
| BenchLM | 59.1 | Overall Score | 72.1 |
| Bullshit Benchmark | 36.4 | BS Detection Rate (%) | 71.9 |
| SWE-rebench | 51.35 | Resolved (%) | 70 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex · How the rankings work · Data refreshed daily, snapshot 2026-07-22.