GPT-5.2 Codex — benchmark results

OpenAI's coding-specialized GPT-5.2 Codex model for software tasks. Provider: OpenAI. Released 2026-01-14. Access: API.

Unified ELO 1760 ± 20, rank #180 of 1776 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA-LCR75.7Score (self-reported)100
Vals AI LiveCodeBench87.99Accuracy (%)98.4
CritPt8.7Accuracy (self-reported)93.3
SnakeBench32.4TrueSkill Rating91.6
AI Chess Leaderboard (Reasoning)1303Elo90
AISI Cyber CTF72Success Rate (%)89.3
AI Chess Leaderboard (Continuation)1151Elo87.2
Terminal-Bench 2.066.5Accuracy (%)86.2
Epoch AI - Apex Agents27.6Score72.9
BenchLM59.1Overall Score72.1
Bullshit Benchmark36.4BS Detection Rate (%)71.9
SWE-rebench51.35Resolved (%)70

Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex · How the rankings work · Data refreshed daily, snapshot 2026-07-22.