GPT-5.3 Codex (High) — benchmark results
GPT-5.3 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2026-02-05. Access: API.
Unified ELO 1879 ± 17, rank #81 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| APEX v1 Investment Banking | 65 | Score (%) | 100 |
| SlopCodeBench | 26.02 | Isolated Solved (%) | 93.8 |
| Epoch AI - ECI | 155.57 | ECI Score | 88.1 |
| APEX v1 Consulting | 65.5 | Score (%) | 83.3 |
| APEX-Agents | 46.9 | Mean Score (ReAct) (self-reported) | 79.5 |
| Epoch AI - Apex Agents | 31.7 | Score | 78.1 |
| PrinzBench | 52 | Score (x/99) | 69.6 |
| BinaryAudit | 82.46 | Avg Success Rate (%) | 68 |
| InferenceBench | 5.49 | Speedup Score | 63.6 |
| LiveBench | 73.18 | LiveBench average (self-reported) | 47.6 |
| APEX v1 Big Law | 69.8 | Score (%) | 33.3 |
| APEX v1 Medicine (MD) | 58.1 | Score (%) | 25 |
Interactive version: theaggregate.ai/model?slug=gpt-5-3-codex-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.