GPT-5.4 Pro (xHigh) — benchmark results

GPT-5.4 Pro evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-03-06. Access: API.

Unified ELO 2019 ± 27, rank #17 of 1776 rated models, from 65 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EnigmaEval23.82Score (self-reported)100
LLMEval-Logic Formalization Fixed60.2Accuracy (%)100
MathArena - ARXIV February75.78Accuracy (%)100
OpenAI GPT-5.4 Launch - ARC-AGI-1 (Verified)94.5Score (%)100
OpenAI GPT-5.4 Launch - ARC-AGI-2 (Verified)83.3Score (%)100
OpenAI GPT-5.4 Launch - BrowseComp89.3Score (%)100
OpenAI GPT-5.4 Launch - FinanceAgent v1.161.5Score (%)100
OpenAI GPT-5.4 Launch - Frontier Science Research36.7Score (%)100
OpenAI GPT-5.4 Launch - FrontierMath Tier 1-350Score (%)100
OpenAI GPT-5.4 Launch - FrontierMath Tier 438Score (%)100
OpenAI GPT-5.4 Launch - GPQA Diamond94.4Score (%)100
OpenAI GPT-5.4 Launch - Humanity's Last Exam (no tools)42.7Score (%)100

Interactive version: theaggregate.ai/model?slug=gpt-5-4-pro-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.