GPT-5.4 (2026-03-05) (Medium): benchmark results

Provider: OpenAI. Released 2026-03-05. Access: API.

Unified ELO 1741 ± 15, rank #255 of 2131 rated models, from 53 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EgoGapBench66.1Accuracy (%; zero-shot multiple choice on 1,000 single-image100
EgoGapBench - Direct Actions50.9Accuracy (%; zero-shot multiple choice on 650 direct-action 100
EgoGapBench - Four Options65.7Accuracy (%; zero-shot multiple choice on 382 four-option si100
EgoGapBench - Indirect Actions94.3Accuracy (%; zero-shot multiple choice on 350 indirect-actio100
EgoGapBench - Three Options66.3Accuracy (%; zero-shot multiple choice on 618 three-option s100
Swallow - English MT-Bench - Extraction85.2Judge Score (normalized, %)100
Swallow - English MT-Bench - Roleplay86.8Judge Score (normalized, %)100
Swallow - English MT-Bench - Writing80.9Judge Score (normalized, %)100
Swallow - Japanese MT-Bench - Average84.4Judge Score (normalized, %)100
Swallow - Japanese MT-Bench - Extraction78.7Judge Score (normalized, %)100
Swallow - Japanese MT-Bench - Writing80.4Judge Score (normalized, %)100
Swallow - Post-trained English - AIME95.8Accuracy (%)100

Interactive version: theaggregate.ai/model?slug=gpt-5-4-2026-03-05-medium · How It Works · Data refreshed daily, snapshot 2026-10-09.