GPT-5.4 (Medium) — benchmark results
GPT-5.4 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2026-03-06. Access: API.
Unified ELO 1930 ± 28, rank #48 of 1776 rated models, from 39 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SMDD-Bench | 40.2 | Pass Rate (%) | 100 |
| AI for Education Visual Maths | 86.5 | Accuracy (%) | 98.3 |
| AI for Education Visual Maths - Geometry | 86.92 | Accuracy (%) | 98.3 |
| AI for Education Visual Maths - Measurement | 97.3 | Accuracy (%) | 98.3 |
| AI for Education Visual Maths - Statistics and Probability | 71.43 | Accuracy (%) | 98.3 |
| AI for Education Pedagogy - Maths | 92.06 | Accuracy (%) | 97.5 |
| AI for Education Pedagogy - Social studies | 90 | Accuracy (%) | 97.5 |
| LLM Chess (Saplin) | 1252.9 | ELO | 97.1 |
| AI for Education Visual Reasoning - reasoning by analogy | 81.4 | Accuracy (%) | 96.8 |
| AI for Education Visual Reasoning - pattern completion (linear) | 87.2 | Accuracy (%) | 96 |
| AI for Education Visual Maths - Algebra | 96.15 | Accuracy (%) | 95.8 |
| AI for Education SEND | 83.94 | Accuracy (%) | 94.1 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.