GPT-5.1 Chat: benchmark results
Provider: OpenAI. Released 2025-11-12. Access: API.
Unified ELO 1717 ± 31, rank #295 of 2656 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchTable - STEM | 77.6 | Weighted Score (%) | 89.2 |
| BenchTable - Reasoning | 67.9 | Weighted Score (%) | 83.4 |
| BenchTable | 65.6 | Total Score (%) | 79.4 |
| Korean CSAT 2026 (Easy Mode) - English | 97 | Points (out of 100) | 66 |
| BenchTable - Utility | 60.1 | Weighted Score (%) | 62.8 |
| BenchTable - Tech | 61.7 | Weighted Score (%) | 59.4 |
| SnakeBench | 23.3 | TrueSkill Rating | 54.7 |
| Korean CSAT 2026 (Easy Mode) - Total | 372.5 | Points (out of 450) | 33.7 |
| Korean CSAT 2026 (Easy Mode) - Physics I | 23 | Points (out of 50) | 31.6 |
| Korean CSAT 2026 (Easy Mode) - Mathematics | 88 | Points (out of 100) | 26.9 |
| Korean CSAT 2026 (Easy Mode) - Society and Culture | 30 | Points (out of 50) | 24.8 |
| Korean CSAT 2026 (Easy Mode) - Life Science I | 30 | Points (out of 50) | 23.8 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-chat · How It Works · Data refreshed daily, snapshot 2026-09-19.