GPT-5.4 (OpenClaw): benchmark results

Provider: OpenAI. Released 2026-03-05. Access: API.

Unified ELO 1844 ± 26, rank #60 of 1605 rated models, from 21 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HarnessAudit-Bench - Runtime Robustness0.68Perturbation stability under tool or runtime errors and nois100
PseudoBench72.6Pseudoscientific hazard (%; mean of report quality, pseudosc100
PseudoBench - Persuasiveness62Pseudoscientific persuasiveness (%; misuse of terminology, a100
HarnessAudit-Bench0.32Overall harness safety score (0-1; task-level safety adheren83.3
Herculean - Auditing (OpenClaw)66.15Accuracy (%; share of the 65 SEC XBRL audit instances where 83.3
PseudoBench - Pseudoscience Alignment70.9Pseudoscience alignment (%; how far the report keeps, uses a83.3
MacAgentBench (OpenClaw) - Multi-App45Pass@1 (%; share of the 140 Multi-App tasks solved on a sing80
HarnessAudit-Bench - Ambiguous Goal Stability0.35Perturbation stability under ambiguous or underspecified goa77.8
HarnessAudit-Bench - Resource Safety Adherence0.39Safety adherence rate, resource channel (0-1; 1 minus the se77.8
WildClawBench50.3Overall Score (self-reported)72.4
MacAgentBench (OpenClaw)60.7Pass@1 (%; share of the 676 macOS desktop tasks in 25 applic60
MacAgentBench (OpenClaw) - Development61.4Pass@1 (%; share of the 44 Development tasks solved on a sin60

Interactive version: theaggregate.ai/model?slug=gpt-5-4-openclaw · How It Works · Data refreshed daily, snapshot 2026-09-27.