GPT-5.4 Nano (2026-03-17): benchmark results
Provider: OpenAI. Released 2026-03-17. Access: API.
Unified ELO 1691 ± 13, rank #215 of 1639 rated models, from 503 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RoboJailBench | 97.84 | Security-utility harmonic mean (%) of the rejection rate on | 100 |
| RoboJailBench - Utility Rate | 99.49 | Utility rate (%), the share of benign goals accepted, over t | 100 |
| EuroEval Bosnian NLU - MMS BS | 53.99 | Sentiment classification Score (%) | 99.5 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 53.49 | Sentiment classification Score (%) | 99.3 |
| Vectara Hallucination Leaderboard | 96.9 | Factual Consistency Rate (%) | 99.1 |
| Horangi 4 - GLP - Mathematical Reasoning | 98 | Score (%) | 98.6 |
| EuroEval Finnish NLU - ScaLA FI | 51.4 | Linguistic acceptability Score (%) | 98.4 |
| EuroEval Polish Summarization - PSC | 25.44 | Score (%) | 97.3 |
| EuroEval Spanish NLU - ScaLA ES | 55.08 | Linguistic acceptability Score (%) | 97.1 |
| EuroEval Portuguese NLU - ScaLA PT | 64.73 | Linguistic acceptability Score (%) | 97 |
| EuroEval Slovene Knowledge | 78.4 | Knowledge Average Score (%) | 96.8 |
| EuroEval Dutch Knowledge | 83.32 | Knowledge Average Score (%) | 96.2 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-nano-2026-03-17 · How It Works · Data refreshed daily, snapshot 2026-10-09.