GPT-5 (Low): benchmark results
GPT-5 evaluated at the low reasoning-effort setting. Provider: OpenAI. Released 2025-08-07. Access: API.
Unified ELO 1642 ± 1, rank #273 of 1761 rated models, from 75 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - Dart | 44 | Accuracy (%) | 97.8 |
| FlagEval VQA - Geographic Reasoning | 70.2 | Score | 97.8 |
| UGI - Natural Intelligence | 67.2 | NatInt Score | 96.7 |
| AA Omniscience - Software Engineering (SWE) - Java | 37 | Accuracy (%) | 96.4 |
| AA Omniscience - Software Engineering (SWE) - PHP | 54 | Accuracy (%) | 95.7 |
| FlagEval VQA - Chart Understanding | 60.3 | Score | 95.7 |
| FlagEval VQA - Subject Knowledge | 71.6 | Score | 95.7 |
| AA Omniscience - Software Engineering (SWE) - Julia | 36 | Accuracy (%) | 94.5 |
| AA Omniscience - Software Engineering (SWE) - Python | 44 | Accuracy (%) | 94.2 |
| AA Omniscience - Software Engineering (SWE) - R | 38 | Accuracy (%) | 93.7 |
| AA MMLU-Pro | 85.98 | Accuracy (%) | 93.4 |
| GAIA2 - Ambiguity | 39.6 | Score (%) | 93.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-low · How It Works · Data refreshed daily, snapshot 2026-09-05.