GPT-5.4 Nano (2026-03-17): benchmark results

Provider: OpenAI. Released 2026-03-17. Access: API.

Unified ELO 1691 ± 13, rank #215 of 1639 rated models, from 503 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RoboJailBench97.84Security-utility harmonic mean (%) of the rejection rate on 100
RoboJailBench - Utility Rate99.49Utility rate (%), the share of benign goals accepted, over t100
EuroEval Bosnian NLU - MMS BS53.99Sentiment classification Score (%)99.5
EuroEval Spanish NLU - Sentiment Headlines ES53.49Sentiment classification Score (%)99.3
Vectara Hallucination Leaderboard96.9Factual Consistency Rate (%)99.1
Horangi 4 - GLP - Mathematical Reasoning98Score (%)98.6
EuroEval Finnish NLU - ScaLA FI51.4Linguistic acceptability Score (%)98.4
EuroEval Polish Summarization - PSC25.44Score (%)97.3
EuroEval Spanish NLU - ScaLA ES55.08Linguistic acceptability Score (%)97.1
EuroEval Portuguese NLU - ScaLA PT64.73Linguistic acceptability Score (%)97
EuroEval Slovene Knowledge78.4Knowledge Average Score (%)96.8
EuroEval Dutch Knowledge83.32Knowledge Average Score (%)96.2

Interactive version: theaggregate.ai/model?slug=gpt-5-4-nano-2026-03-17 · How It Works · Data refreshed daily, snapshot 2026-10-09.