GPT-5.2 Codex (High): benchmark results
GPT-5.2 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-12-18. Access: API.
Unified ELO 1747 ± 26, rank #235 of 2088 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BinaryAudit | 88.41 | Avg Success Rate (%) | 88 |
| SlopCodeBench | 21.94 | Isolated Solved (%) | 81.2 |
| APEX-Agents | 42.2 | Mean Score (ReAct) (self-reported) | 79.5 |
| WeaveBench - Data Analysis and Visualization | 15.4 | PassRate (%) on the Data Analysis and Visualization domain t | 72.2 |
| Sonar LLM Leaderboard - Java - Pass Rate | 80.35 | Passing tests (%) | 68.6 |
| WeaveBench - DevOps | 16.7 | PassRate (%) on the DevOps domain tasks of 114 long-horizon | 61.1 |
| Sonar LLM Leaderboard - Java - Code Smell Density | 17.73 | Code smells per 1,000 lines of code (lower is better) | 58.6 |
| Sonar LLM Leaderboard - Java - Bug Density | 0.75 | Bugs per 1,000 lines of code (lower is better) | 57.9 |
| WeaveBench - Desktop | 5.6 | PassRate (%) on the Desktop domain tasks of 114 long-horizon | 55.6 |
| WeaveBench - Document | 11.8 | PassRate (%) on the Document domain tasks of 114 long-horizo | 55.6 |
| WeaveBench - Overall Score | 32.1 | Mean per-task score (0-100) on 114 long-horizon real-world c | 55.6 |
| LitXBench (Coding Agents) - Materials | 0.95 | Material F1 (0-1), whether the right set of materials is ext | 50 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex-high · How It Works · Data refreshed daily, snapshot 2026-10-07.