GPT-5.4 Mini (2026-03-17) (xHigh): benchmark results
Provider: OpenAI. Released 2026-03-17. Access: API.
Unified ELO 1719 ± 17, rank #301 of 2131 rated models, from 57 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Horangi 4 - SWE-bench Verified (80-task subset) | 67.5 | Resolved (%) | 95.7 |
| Horangi 4 - BigCodeBench (100-task subset) | 59 | Pass rate (%) | 95.2 |
| Horangi 4 - HRM8K | 96 | Accuracy (%) | 93.8 |
| Horangi 4 - GLP - Coding | 75.17 | Score (%) | 93.3 |
| Horangi 4 - IFEval-Ko | 92.5 | Instruction-following accuracy (%) | 93.3 |
| Horangi 4 - GLP - Mathematical Reasoning | 96.33 | Score (%) | 90.4 |
| Horangi 4 - KoBBQ | 95 | Accuracy (%) | 88.5 |
| Horangi 4 - Ko-AIME 2025 | 96.67 | Accuracy (%) | 87 |
| Horangi 4 - HumanEval (100-problem subset) | 99 | Pass rate (%) | 86.5 |
| Horangi 4 - Ko-Moral | 82 | Accuracy (%) | 85.6 |
| Horangi 4 - Ko-HalluLens (WikiQA) | 44 | Correct answer rate (%) | 84.6 |
| Vals AI SAGE | 50.81 | Accuracy (%) | 84.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-mini-2026-03-17-xhigh · How It Works · Data refreshed daily, snapshot 2026-10-09.