GPT-5.4 Nano (Low): benchmark results
Provider: OpenAI. Released 2026-03-17. Access: API.
Unified ELO 1539 ± 30, rank #905 of 2131 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FrontierMath - Tiers 1-3 (v2) | 20.35 | Accuracy (%, 285 private v2 problems) | 31.3 |
| MageBench S2 | 1584 | Combined 1v1 Elo | 20 |
| ARC-AGI-2 | 1.53 | Accuracy (%) | 18.2 |
| ARC-AGI-1 | 18.33 | Accuracy (%) | 15 |
| ObviousBench | 52.78 | Answer pass³ (%) | 13.2 |
| T1-Bench - Tool Call F1 | 73.66 | Tool-name F1 (%) of the assistant's tool calls against the g | 9.1 |
| Context Arena | 21.87 | Average Score (%) | 7.3 |
| o11y-bench - Pass@3 | 50.79 | Tasks passed on at least one of three attempts, Pass@3 (%) | 3.5 |
| o11y-bench - Pass^3 | 19.05 | Tasks passed on all three attempts, Pass^3 (%) | 1.8 |
| Equation-Suffix Prediction (Kimi K2.6 Scorer) | 0.05 | Likelihood lift (mean clipLL2 per target token over the same | 0 |
| Equation-Suffix Prediction (Qwen3-8B Scorer) | 0.08 | Likelihood lift (mean clipLL2 per target token over the same | 0 |
| T1-Bench | 20.38 | Pass@3 (%): share of conversations solved in at least one of | 0 |
Interactive version: theaggregate.ai/model?slug=gpt-5-4-nano-low · How It Works · Data refreshed daily, snapshot 2026-10-09.