GPT-5.4 (Codex): benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1765 ± 16, rank #132 of 1605 rated models, from 38 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CyberChainBench - Exploit | 43.3 | Exploit reward (%; min(agent profit / reference profit, 1) f | 88.9 |
| Herculean - Auditing (Codex) | 63.08 | Accuracy (%; share of the 65 SEC XBRL audit instances where | 83.3 |
| CliniCARE-Bench - Defect-Free Accuracy | 65.6 | Correct verdicts whose report has no judged defect (%) | 76.7 |
| DrawAI-Bench - Editability - Formula | 97.8 | Editability score (0-100; rules and VLM rubric, strongest ob | 75 |
| CliniCARE-Bench - Expected Calibration Error | 0.15 | Expected calibration error (0-1; lower is better) | 73.3 |
| CliniCARE-Bench - Over-Commitment Rate | 36.7 | Definitive verdicts on cases whose reference requires deferr | 73.3 |
| CliniCARE-Bench - Process Rubric Score | 76.1 | Process rubric score (%; GPT-5.5 judge) | 73.3 |
| CyberChainBench - Detect | 32.5 | Detection reward (%; 1 when both the root-cause vulnerabilit | 66.7 |
| CyberChainBench - Patch | 19.1 | Patch reward (%; 1 when the patched implementation makes the | 66.7 |
| DrawAI-Bench - Fidelity - Shape | 92 | Fidelity score (0-100; rules and VLM rubric, strongest obser | 66.7 |
| PatchEval-Verified | 78.7 | Pass@1 (%) | 62.5 |
| SkillSafetyBench - Memory, Recovery, Audit and Persistence | 46.2 | Attack success rate (%; 26 cases of unsafe state that persis | 62.5 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-codex · How It Works · Data refreshed daily, snapshot 2026-09-26.