GPT-5.3 Codex (xHigh): benchmark results
GPT-5.3 Codex evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-02-05. Access: API.
Unified ELO 1701 ± 1, rank #103 of 1761 rated models, from 33 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ALE-Bench | 1655.22 | Performance (Self-Refine x1) (self-reported) | 98.9 |
| AA Long Context Reasoning | 83.33 | Accuracy (%) | 97.9 |
| AA Omniscience - Software Engineering (SWE) | 85.2 | Accuracy (%) | 97 |
| AA Terminal-Bench Hard | 53.03 | Accuracy (%) | 96.2 |
| AA Omniscience - Health | 48.1 | Accuracy (%) | 94.7 |
| PM-LLM-Benchmark | 37.3 | Score | 94 |
| AA Humanity's Last Exam | 42.49 | Accuracy (%) | 93.8 |
| AA-Omniscience Accuracy | 52.88 | Accuracy (%) | 93.4 |
| AA IFBench | 75.37 | Accuracy (%) | 92.7 |
| AA GPQA Diamond | 91.52 | Accuracy (%) | 92.6 |
| AA Omniscience - Business | 42.1 | Accuracy (%) | 92 |
| AA-LCR | 78.33 | Accuracy (self-reported) | 92 |
Interactive version: theaggregate.ai/model?slug=gpt-5-3-codex-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.