GPT-5.3 Codex (xHigh): benchmark results

GPT-5.3 Codex evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-02-05. Access: API.

Unified ELO 1701 ± 1, rank #103 of 1761 rated models, from 33 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ALE-Bench1655.22Performance (Self-Refine x1) (self-reported)98.9
AA Long Context Reasoning83.33Accuracy (%)97.9
AA Omniscience - Software Engineering (SWE)85.2Accuracy (%)97
AA Terminal-Bench Hard53.03Accuracy (%)96.2
AA Omniscience - Health48.1Accuracy (%)94.7
PM-LLM-Benchmark37.3Score94
AA Humanity's Last Exam42.49Accuracy (%)93.8
AA-Omniscience Accuracy52.88Accuracy (%)93.4
AA IFBench75.37Accuracy (%)92.7
AA GPQA Diamond91.52Accuracy (%)92.6
AA Omniscience - Business42.1Accuracy (%)92
AA-LCR78.33Accuracy (self-reported)92

Interactive version: theaggregate.ai/model?slug=gpt-5-3-codex-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.