GPT-5.4 (Codex CLI): benchmark results

Provider: OpenAI. Released 2026-03-05. Access: API.

Unified ELO 1750 ± 22, rank #153 of 1605 rated models, from 24 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SkillEvolBench (Curated-Revision-Always) - Deployment Success40Frozen deployment success rate (%; ESR, share of the 90 cont100
HarnessAudit-Bench - Indirect Injection Stability0.24Perturbation stability under indirect prompt injection (0-1;94.4
ProcCtrlBench0.73ProcCtrlBench score PB (0-1; calibrated process quality over90
HarnessAudit-Bench - Task Completion0.76Task completion rate (0-1; weighted hidden completion checkp88.9
SkillEvolBench (Curated-Revision) - Deployment Success38.9Frozen deployment success rate (%; ESR, share of the 90 cont88.9
WildClawBench56.8Overall Score (self-reported)86.6
SkillEvolBench (No-Skill) - Acquisition Success43.3Acquisition success rate (%; LSR, share of the 90 canonical,83.3
ResearchClawBench18.4Overall (self-reported)82.6
HarnessAudit-Bench - Runtime Robustness0.57Perturbation stability under tool or runtime errors and nois72.2
PseudoBench - Persuasiveness74.2Pseudoscientific persuasiveness (%; misuse of terminology, a66.7
PseudoBench - Pseudoscience Alignment77.6Pseudoscience alignment (%; how far the report keeps, uses a66.7
CyberGym66.3Success Rate (self-reported)54

Interactive version: theaggregate.ai/model?slug=gpt-5-4-codex-cli · How It Works · Data refreshed daily, snapshot 2026-09-27.