GPT-5.3 Codex (xHigh) — benchmark results
GPT-5.3 Codex evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-02-05. Access: API.
Unified ELO 1929 ± 28, rank #49 of 1776 rated models, from 44 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - JavaScript | 90.91 | Accuracy (%) | 100 |
| AA Omniscience - Software Engineering (SWE) - Dart | 80 | Accuracy (%) | 99.8 |
| AA Omniscience - Software Engineering (SWE) - Java | 73 | Accuracy (%) | 99.8 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 90 | Accuracy (%) | 99.8 |
| AA Omniscience - Software Engineering (SWE) - Julia | 84 | Accuracy (%) | 99.4 |
| AA Omniscience - Software Engineering (SWE) - Rust | 86 | Accuracy (%) | 99.2 |
| AA Omniscience - Software Engineering (SWE) | 83.8 | Accuracy (%) | 99 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 90 | Accuracy (%) | 98.8 |
| AA Omniscience - Software Engineering (SWE) - Python | 87.5 | Accuracy (%) | 98.7 |
| ALE-Bench | 1655.22 | Performance (Self-Refine x1) (self-reported) | 98.7 |
| AA Omniscience - Software Engineering (SWE) - HTML | 86 | Accuracy (%) | 98.5 |
| AA Long Context Reasoning | 74 | Accuracy (%) | 98.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-3-codex-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.