GPT-5.4 (High): benchmark results

GPT-5.4 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2026-03-06. Access: API.

Unified ELO 1702 ± 1, rank #97 of 1761 rated models, from 90 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CubeBench66.7Success Rate (%)100
LLM2014 Logic 2026-0378.85Median Score100
Persuasion (Lechmazur)1.71Average Persuasion Strength100
UGI - Natural Intelligence71.25NatInt Score98.1
UGI - Writing68.82Writing Score97.8
ALE-Bench1607Performance (Self-Refine x1) (self-reported)97.7
LLM2014 Logic 2026-0479.56Median Score97.5
Buyout Game (Lechmazur)1960.6Bradley-Terry Rating97.1
LLM Chess (Saplin)1322.3ELO96.9
Chatbot Arena (Text - Multi-Turn)1494Arena Score96.5
SkateBench81.54Success Rate (%)96.3
Chatbot Arena (Text - Math)1494Arena Score96.1

Interactive version: theaggregate.ai/model?slug=gpt-5-4-high · How It Works · Data refreshed daily, snapshot 2026-09-05.