GPT-5.1 (Thinking): benchmark results
GPT-5.1 evaluated with thinking enabled. Provider: OpenAI. Released 2025-11-12. Access: API.
Unified ELO 1667 ± 1, rank #194 of 1761 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SuperGPQA | 66.35 | Accuracy (%) | 97.1 |
| ALE-Bench | 1192.15 | Performance (Self-Refine x1) (self-reported) | 85.2 |
| SEAL - MultiChallenge | 63.41 | Score | 82.8 |
| SEAL - Humanity's Last Exam (Text Only) | 24.65 | Score | 81.7 |
| LLM Stats Score | 36.78 | LLM Stats Score (conservative rating) | 80.6 |
| SEAL - MultiNRC | 49 | Score | 77.9 |
| SEAL - MASK | 86.33 | Score | 77.3 |
| SEAL - TutorBench | 54.09 | Score | 76.9 |
| SEAL - Professional Reasoning Benchmark - Legal | 49.33 | Score | 74.2 |
| EnigmaEval | 11.23 | Score (self-reported) | 74 |
| SEAL - Humanity's Last Exam | 23.68 | Score | 74 |
| SEAL - Professional Reasoning Benchmark - Finance | 48.01 | Score | 71 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.