GPT-5.1 (Thinking) — benchmark results
GPT-5.1 evaluated with thinking enabled. Provider: OpenAI. Released 2025-11-12. Access: API.
Unified ELO 1783 ± 22, rank #159 of 1776 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SuperGPQA | 66.35 | Accuracy (%) | 97.1 |
| ZeroEval GPQA Diamond | 88.1 | GPQA Diamond Score | 86.5 |
| SEAL - EnigmaEval | 11.23 | Score | 83.1 |
| SEAL - MultiChallenge | 63.41 | Score | 82.8 |
| SEAL - Professional Reasoning Benchmark - Legal | 49.33 | Score | 82.1 |
| SEAL - Professional Reasoning Benchmark - Finance | 48.01 | Score | 78.6 |
| SEAL - MultiNRC | 49 | Score | 77.9 |
| SEAL - MASK | 86.33 | Score | 77.3 |
| SEAL - TutorBench | 54.09 | Score | 76.9 |
| SEAL - Humanity's Last Exam | 23.68 | Score | 75.5 |
| SEAL - Fortress | 25.72 | Score | 49.1 |
| SEAL - VISTA | 43.82 | Score | 46.8 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.