GPT-5.4 (Codex CLI): benchmark results
Provider: OpenAI. Released 2026-03-05. Access: API.
Unified ELO 1750 ± 22, rank #153 of 1605 rated models, from 24 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SkillEvolBench (Curated-Revision-Always) - Deployment Success | 40 | Frozen deployment success rate (%; ESR, share of the 90 cont | 100 |
| HarnessAudit-Bench - Indirect Injection Stability | 0.24 | Perturbation stability under indirect prompt injection (0-1; | 94.4 |
| ProcCtrlBench | 0.73 | ProcCtrlBench score PB (0-1; calibrated process quality over | 90 |
| HarnessAudit-Bench - Task Completion | 0.76 | Task completion rate (0-1; weighted hidden completion checkp | 88.9 |
| SkillEvolBench (Curated-Revision) - Deployment Success | 38.9 | Frozen deployment success rate (%; ESR, share of the 90 cont | 88.9 |
| WildClawBench | 56.8 | Overall Score (self-reported) | 86.6 |
| SkillEvolBench (No-Skill) - Acquisition Success | 43.3 | Acquisition success rate (%; LSR, share of the 90 canonical, | 83.3 |
| ResearchClawBench | 18.4 | Overall (self-reported) | 82.6 |
| HarnessAudit-Bench - Runtime Robustness | 0.57 | Perturbation stability under tool or runtime errors and nois | 72.2 |
| PseudoBench - Persuasiveness | 74.2 | Pseudoscientific persuasiveness (%; misuse of terminology, a | 66.7 |
| PseudoBench - Pseudoscience Alignment | 77.6 | Pseudoscience alignment (%; how far the report keeps, uses a | 66.7 |
| CyberGym | 66.3 | Success Rate (self-reported) | 54 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-codex-cli · How It Works · Data refreshed daily, snapshot 2026-09-27.