GPT-5.4 (2026-03-05) (Medium): benchmark results
Provider: OpenAI. Released 2026-03-05. Access: API.
Unified ELO 1741 ± 15, rank #255 of 2131 rated models, from 53 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EgoGapBench | 66.1 | Accuracy (%; zero-shot multiple choice on 1,000 single-image | 100 |
| EgoGapBench - Direct Actions | 50.9 | Accuracy (%; zero-shot multiple choice on 650 direct-action | 100 |
| EgoGapBench - Four Options | 65.7 | Accuracy (%; zero-shot multiple choice on 382 four-option si | 100 |
| EgoGapBench - Indirect Actions | 94.3 | Accuracy (%; zero-shot multiple choice on 350 indirect-actio | 100 |
| EgoGapBench - Three Options | 66.3 | Accuracy (%; zero-shot multiple choice on 618 three-option s | 100 |
| Swallow - English MT-Bench - Extraction | 85.2 | Judge Score (normalized, %) | 100 |
| Swallow - English MT-Bench - Roleplay | 86.8 | Judge Score (normalized, %) | 100 |
| Swallow - English MT-Bench - Writing | 80.9 | Judge Score (normalized, %) | 100 |
| Swallow - Japanese MT-Bench - Average | 84.4 | Judge Score (normalized, %) | 100 |
| Swallow - Japanese MT-Bench - Extraction | 78.7 | Judge Score (normalized, %) | 100 |
| Swallow - Japanese MT-Bench - Writing | 80.4 | Judge Score (normalized, %) | 100 |
| Swallow - Post-trained English - AIME | 95.8 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-2026-03-05-medium · How It Works · Data refreshed daily, snapshot 2026-10-09.