GPT-5.4 Mini (2026-03-17) (xHigh): benchmark results

Provider: OpenAI. Released 2026-03-17. Access: API.

Unified ELO 1719 ± 17, rank #301 of 2131 rated models, from 57 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Horangi 4 - SWE-bench Verified (80-task subset)67.5Resolved (%)95.7
Horangi 4 - BigCodeBench (100-task subset)59Pass rate (%)95.2
Horangi 4 - HRM8K96Accuracy (%)93.8
Horangi 4 - GLP - Coding75.17Score (%)93.3
Horangi 4 - IFEval-Ko92.5Instruction-following accuracy (%)93.3
Horangi 4 - GLP - Mathematical Reasoning96.33Score (%)90.4
Horangi 4 - KoBBQ95Accuracy (%)88.5
Horangi 4 - Ko-AIME 202596.67Accuracy (%)87
Horangi 4 - HumanEval (100-problem subset)99Pass rate (%)86.5
Horangi 4 - Ko-Moral82Accuracy (%)85.6
Horangi 4 - Ko-HalluLens (WikiQA)44Correct answer rate (%)84.6
Vals AI SAGE50.81Accuracy (%)84.3

Interactive version: theaggregate.ai/model?slug=gpt-5-4-mini-2026-03-17-xhigh · How It Works · Data refreshed daily, snapshot 2026-10-09.