GPT-5 Nano (2025-08-07) (Medium): benchmark results
Provider: OpenAI. Released 2025-08-07. Access: API.
Unified ELO 1550 ± 32, rank #863 of 2131 rated models, from 21 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BeQu - Experiment 1 - Entailment Precision | 88 | Entailment Precision (%) | 100 |
| MATH Level 5 | 95.24 | Accuracy (%) | 87.9 |
| OTIS Mock AIME 2024-25 | 74.17 | Accuracy (%) | 57.6 |
| Epoch AI - GPQA Diamond | 67.42 | Accuracy (%) | 40.9 |
| CommunityFact (Closed-Input) - Finance (English) | 57.77 | Macro-F1 (%) over True and False claims on the temporally he | 35.7 |
| CommunityFact (Closed-Input) - Politics (English) | 55.47 | Macro-F1 (%) over True and False claims on the temporally he | 35.7 |
| CommunityFact (Closed-Input) - Finance (French) | 52.56 | Macro-F1 (%) over True and False claims on the temporally he | 28.6 |
| CommunityFact (Closed-Input) - Finance (Japanese) | 61.8 | Macro-F1 (%) over True and False claims on the temporally he | 28.6 |
| CommunityFact (Closed-Input) - Finance (Portuguese) | 53.14 | Macro-F1 (%) over True and False claims on the temporally he | 28.6 |
| CommunityFact (Closed-Input) - Finance (Spanish) | 51.29 | Macro-F1 (%) over True and False claims on the temporally he | 28.6 |
| CommunityFact (Closed-Input) - Politics (Portuguese) | 55.35 | Macro-F1 (%) over True and False claims on the temporally he | 28.6 |
| VPCT | 35.4 | Accuracy (%) | 25.6 |
Interactive version: theaggregate.ai/model?slug=gpt-5-nano-2025-08-07-medium · How It Works · Data refreshed daily, snapshot 2026-10-09.