GPT-5.4 (Non-reasoning) — benchmark results
GPT-5.4 evaluated with reasoning disabled. Provider: OpenAI. Released 2026-03-06. Access: API.
Unified ELO 1723 ± 32, rank #234 of 1776 rated models, from 45 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - Julia | 72 | Accuracy (%) | 97.1 |
| AA Omniscience - Software Engineering (SWE) - Python | 74.5 | Accuracy (%) | 94.5 |
| AA Omniscience - Software Engineering (SWE) - Rust | 78 | Accuracy (%) | 93.9 |
| AA Omniscience - Software Engineering (SWE) - Dart | 58 | Accuracy (%) | 93.8 |
| UGI - Natural Intelligence | 57.24 | NatInt Score | 93.5 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 70 | Accuracy (%) | 93.2 |
| UGI - Writing | 57.99 | Writing Score | 92.7 |
| AA Omniscience - Software Engineering (SWE) - R | 54 | Accuracy (%) | 91.9 |
| AA Omniscience - Software Engineering (SWE) | 64.6 | Accuracy (%) | 91.5 |
| AA Omniscience - Software Engineering (SWE) - C | 75 | Accuracy (%) | 91.3 |
| Epoch AI - ECI | 156.12 | ECI Score | 91.2 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 58 | Accuracy (%) | 91 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.