GPT-5.2 Codex (High): benchmark results

GPT-5.2 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-12-18. Access: API.

Unified ELO 1747 ± 26, rank #235 of 2088 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BinaryAudit88.41Avg Success Rate (%)88
SlopCodeBench21.94Isolated Solved (%)81.2
APEX-Agents42.2Mean Score (ReAct) (self-reported)79.5
WeaveBench - Data Analysis and Visualization15.4PassRate (%) on the Data Analysis and Visualization domain t72.2
Sonar LLM Leaderboard - Java - Pass Rate80.35Passing tests (%)68.6
WeaveBench - DevOps16.7PassRate (%) on the DevOps domain tasks of 114 long-horizon 61.1
Sonar LLM Leaderboard - Java - Code Smell Density17.73Code smells per 1,000 lines of code (lower is better)58.6
Sonar LLM Leaderboard - Java - Bug Density0.75Bugs per 1,000 lines of code (lower is better)57.9
WeaveBench - Desktop5.6PassRate (%) on the Desktop domain tasks of 114 long-horizon55.6
WeaveBench - Document11.8PassRate (%) on the Document domain tasks of 114 long-horizo55.6
WeaveBench - Overall Score32.1Mean per-task score (0-100) on 114 long-horizon real-world c55.6
LitXBench (Coding Agents) - Materials0.95Material F1 (0-1), whether the right set of materials is ext50

Interactive version: theaggregate.ai/model?slug=gpt-5-2-codex-high · How It Works · Data refreshed daily, snapshot 2026-10-07.