GPT-5.2 Codex (xHigh): benchmark results

GPT-5.2 Codex evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-01-14. Access: API.

Unified ELO 1689 ± 1, rank #134 of 1761 rated models, from 24 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Context-Bench Filesystem93Rubric Score (%)100
AA IFBench77.62Accuracy (%)96.4
BinaryAudit90.67Avg Success Rate (%)96
AA Long Context Reasoning82.33Accuracy (%)95.8
AA-LCR79.33Accuracy (self-reported)95.4
ALE-Bench1299.9Performance (Self-Refine x1) (self-reported)92
AA GPQA Diamond89.9Accuracy (%)89.2
AA Humanity's Last Exam35.73Accuracy (%)87.5
AA TAU-2 Bench92.11Accuracy (%)87.5
AA Omniscience - Health39.6Accuracy (%)86.7
Artificial Analysis Intelligence Index33Intelligence Index85.2
AA Omniscience - Humanities & Social Sciences39.94Accuracy (%)84.8

Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.