GPT-5.2 Codex (High) — benchmark results
GPT-5.2 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2026-01-14. Access: API.
Unified ELO 1877 ± 11, rank #83 of 1776 rated models, from 8 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| APEX v1 Consulting | 66.9 | Score (%) | 100 |
| BinaryAudit | 88.41 | Avg Success Rate (%) | 88 |
| SlopCodeBench | 21.94 | Isolated Solved (%) | 81.2 |
| APEX v1 Investment Banking | 61.8 | Score (%) | 77.8 |
| APEX-Agents | 42.2 | Mean Score (ReAct) (self-reported) | 76.9 |
| APEX v1 Big Law | 73 | Score (%) | 55.6 |
| APEX v1 Medicine (MD) | 59.7 | Score (%) | 50 |
| APEX v1 | 65.3 | Score (%) | 37.5 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.