GPT-5 Mini (2025-08-07) (Medium): benchmark results
Provider: OpenAI. Released 2025-08-07. Access: API.
Unified ELO 1681 ± 16, rank #412 of 2131 rated models, from 47 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - English MT-Bench - Average | 86.9 | Judge Score (normalized, %) | 100 |
| Swallow - Japanese MT-Bench - Coding | 87.9 | Judge Score (normalized, %) | 100 |
| Swallow - Japanese MT-Bench - STEM | 83.8 | Judge Score (normalized, %) | 100 |
| Swallow - English MT-Bench - Coding | 86.1 | Judge Score (normalized, %) | 98.5 |
| Swallow - English MT-Bench - Extraction | 84.5 | Judge Score (normalized, %) | 98.5 |
| Swallow - English MT-Bench - Humanities | 83.4 | Judge Score (normalized, %) | 98.5 |
| Swallow - English MT-Bench - Reasoning | 91 | Judge Score (normalized, %) | 98.5 |
| Swallow - English MT-Bench - STEM | 84.9 | Judge Score (normalized, %) | 98.5 |
| Swallow - English MT-Bench - Writing | 80.5 | Judge Score (normalized, %) | 98.5 |
| Swallow - English MT-Bench - Roleplay | 85.3 | Judge Score (normalized, %) | 97 |
| Swallow - Japanese MT-Bench - Average | 83 | Judge Score (normalized, %) | 97 |
| Swallow - Japanese MT-Bench - Humanities | 80.3 | Judge Score (normalized, %) | 97 |
Interactive version: theaggregate.ai/model?slug=gpt-5-mini-2025-08-07-medium · How It Works · Data refreshed daily, snapshot 2026-10-09.