GPT-5.4 (2026-03-05) (Low): benchmark results
Provider: OpenAI. Released 2026-03-05. Access: API.
Unified ELO 1729 ± 28, rank #277 of 2131 rated models, from 19 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PLCC - Geography | 97 | Accuracy (%) | 92.2 |
| PLCC - Art & Entertainment | 87 | Accuracy (%) | 90.3 |
| PLCC - Culture & Tradition | 93 | Accuracy (%) | 90.3 |
| PLCC - Overall | 90.5 | Mean category accuracy (%) | 90 |
| PLCC - History | 93 | Accuracy (%) | 87.7 |
| EgoGapBench | 58.3 | Accuracy (%; zero-shot multiple choice on 1,000 single-image | 87.5 |
| EgoGapBench - Direct Actions | 40.9 | Accuracy (%; zero-shot multiple choice on 650 direct-action | 87.5 |
| EgoGapBench - Four Options | 57.3 | Accuracy (%; zero-shot multiple choice on 382 four-option si | 87.5 |
| EgoGapBench - Indirect Actions | 90.6 | Accuracy (%; zero-shot multiple choice on 350 indirect-actio | 87.5 |
| EgoGapBench - Three Options | 58.9 | Accuracy (%; zero-shot multiple choice on 618 three-option s | 87.5 |
| PLCC - Grammar | 88 | Accuracy (%) | 86.9 |
| PLCC - Vocabulary | 85 | Accuracy (%) | 84.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-2026-03-05-low · How It Works · Data refreshed daily, snapshot 2026-10-09.