GPT-5.4 (High): benchmark results
GPT-5.4 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2026-03-06. Access: API.
Unified ELO 1702 ± 1, rank #97 of 1761 rated models, from 90 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CubeBench | 66.7 | Success Rate (%) | 100 |
| LLM2014 Logic 2026-03 | 78.85 | Median Score | 100 |
| Persuasion (Lechmazur) | 1.71 | Average Persuasion Strength | 100 |
| UGI - Natural Intelligence | 71.25 | NatInt Score | 98.1 |
| UGI - Writing | 68.82 | Writing Score | 97.8 |
| ALE-Bench | 1607 | Performance (Self-Refine x1) (self-reported) | 97.7 |
| LLM2014 Logic 2026-04 | 79.56 | Median Score | 97.5 |
| Buyout Game (Lechmazur) | 1960.6 | Bradley-Terry Rating | 97.1 |
| LLM Chess (Saplin) | 1322.3 | ELO | 96.9 |
| Chatbot Arena (Text - Multi-Turn) | 1494 | Arena Score | 96.5 |
| SkateBench | 81.54 | Success Rate (%) | 96.3 |
| Chatbot Arena (Text - Math) | 1494 | Arena Score | 96.1 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-high · How It Works · Data refreshed daily, snapshot 2026-09-05.