GPT-5.4 (OpenClaw): benchmark results
Provider: OpenAI. Released 2026-03-05. Access: API.
Unified ELO 1844 ± 26, rank #60 of 1605 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HarnessAudit-Bench - Runtime Robustness | 0.68 | Perturbation stability under tool or runtime errors and nois | 100 |
| PseudoBench | 72.6 | Pseudoscientific hazard (%; mean of report quality, pseudosc | 100 |
| PseudoBench - Persuasiveness | 62 | Pseudoscientific persuasiveness (%; misuse of terminology, a | 100 |
| HarnessAudit-Bench | 0.32 | Overall harness safety score (0-1; task-level safety adheren | 83.3 |
| Herculean - Auditing (OpenClaw) | 66.15 | Accuracy (%; share of the 65 SEC XBRL audit instances where | 83.3 |
| PseudoBench - Pseudoscience Alignment | 70.9 | Pseudoscience alignment (%; how far the report keeps, uses a | 83.3 |
| MacAgentBench (OpenClaw) - Multi-App | 45 | Pass@1 (%; share of the 140 Multi-App tasks solved on a sing | 80 |
| HarnessAudit-Bench - Ambiguous Goal Stability | 0.35 | Perturbation stability under ambiguous or underspecified goa | 77.8 |
| HarnessAudit-Bench - Resource Safety Adherence | 0.39 | Safety adherence rate, resource channel (0-1; 1 minus the se | 77.8 |
| WildClawBench | 50.3 | Overall Score (self-reported) | 72.4 |
| MacAgentBench (OpenClaw) | 60.7 | Pass@1 (%; share of the 676 macOS desktop tasks in 25 applic | 60 |
| MacAgentBench (OpenClaw) - Development | 61.4 | Pass@1 (%; share of the 44 Development tasks solved on a sin | 60 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-openclaw · How It Works · Data refreshed daily, snapshot 2026-09-27.