GPT-5.4 Pro (xHigh): benchmark results

GPT-5.4 Pro evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-03-06. Access: API.

Unified ELO 1729 ± 1, rank #52 of 1761 rated models, from 60 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLMEval-Logic Formalization Fixed60.2Accuracy (%)100
MathArena - ARXIV February75.78Accuracy (%)100
OpenAI GPT-5.4 Launch - ARC-AGI-1 (Verified)94.5Score (%)100
OpenAI GPT-5.4 Launch - ARC-AGI-2 (Verified)83.3Score (%)100
OpenAI GPT-5.4 Launch - BrowseComp89.3Score (%)100
OpenAI GPT-5.4 Launch - FinanceAgent v1.161.5Score (%)100
OpenAI GPT-5.4 Launch - Frontier Science Research36.7Score (%)100
OpenAI GPT-5.4 Launch - FrontierMath Tier 438Score (%)100
OpenAI GPT-5.4 Launch - GPQA Diamond94.4Score (%)100
OpenAI GPT-5.5 Launch - GPQA Diamond94.4Score (%)100
GIM2.16IRT ability (theta)98.9
AA CritPt30Accuracy (%)98.7

Interactive version: theaggregate.ai/model?slug=gpt-5-4-pro-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.