GPT-5.2 Codex (xHigh) — benchmark results

GPT-5.2 Codex evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-01-14. Access: API.

Unified ELO 1882 ± 29, rank #78 of 1776 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Long Context Reasoning75.67Accuracy (%)100
Context-Bench Filesystem93Rubric Score (%)100
AA SciCode54.63Accuracy (%)97.6
AA IFBench77.62Accuracy (%)96.4
BinaryAudit90.67Avg Success Rate (%)96
AA GPQA Diamond89.9Accuracy (%)94.1
AA Omniscience - Health39.1Accuracy (%)92.5
AA Humanity's Last Exam33.46Accuracy (%)92.4
ALE-Bench1299.9Performance (Self-Refine x1) (self-reported)92.3
AA Omniscience - Software Engineering (SWE) - Kotlin60Accuracy (%)92
AA Omniscience - Software Engineering (SWE) - Dart52Accuracy (%)91.4
Artificial Analysis Intelligence Index40.14Intelligence Index90.2

Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.