GPT-5.4 (Medium): benchmark results
GPT-5.4 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2026-03-06. Access: API.
Unified ELO 1703 ± 1, rank #92 of 1761 rated models, from 50 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SMDD-Bench | 40.2 | Pass Rate (%) | 100 |
| AI for Education Pedagogy - Social studies | 90 | Accuracy (%) | 97.7 |
| ALE-Bench | 1520.72 | Performance (Self-Refine x1) (self-reported) | 96.6 |
| AI for Education Visual Maths - Geometry | 86.92 | Accuracy (%) | 95.9 |
| AI for Education Visual Maths - Measurement | 97.3 | Accuracy (%) | 95.9 |
| AI for Education Visual Maths - Statistics and Probability | 71.43 | Accuracy (%) | 95.9 |
| AI for Education Pedagogy - Maths | 92.06 | Accuracy (%) | 95.6 |
| LLM Chess (Saplin) | 1252.9 | ELO | 95 |
| AI for Education Visual Maths | 86.5 | Accuracy (%) | 94.6 |
| AI for Education Visual Maths - Algebra | 96.15 | Accuracy (%) | 94.6 |
| AI for Education Visual Reasoning - pattern completion (linear) | 87.2 | Accuracy (%) | 93.1 |
| AI for Education Visual Reasoning - reasoning by analogy | 81.4 | Accuracy (%) | 92.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.