GPT-5.3 Codex (High): benchmark results
GPT-5.3 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2026-02-05. Access: API.
Unified ELO 1719 ± 1, rank #65 of 1761 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SlopCodeBench | 26.02 | Isolated Solved (%) | 93.8 |
| Vals AI Terminal-Bench 2.0 | 64.05 | Accuracy (%) | 92.4 |
| APEX-Agents | 46.9 | Mean Score (ReAct) (self-reported) | 86.4 |
| PrinzBench | 52 | Score (x/99) | 69.6 |
| BinaryAudit | 82.46 | Avg Success Rate (%) | 68 |
| InferenceBench | 5.49 | Speedup Score | 54.8 |
| LiveBench | 73.18 | LiveBench average (self-reported) | 51.1 |
Interactive version: theaggregate.ai/model?slug=gpt-5-3-codex-high · How It Works · Data refreshed daily, snapshot 2026-09-05.