GPT-5.1 (Thinking): benchmark results

GPT-5.1 evaluated with thinking enabled. Provider: OpenAI. Released 2025-11-12. Access: API.

Unified ELO 1667 ± 1, rank #194 of 1761 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SuperGPQA66.35Accuracy (%)97.1
ALE-Bench1192.15Performance (Self-Refine x1) (self-reported)85.2
SEAL - MultiChallenge63.41Score82.8
SEAL - Humanity's Last Exam (Text Only)24.65Score81.7
LLM Stats Score36.78LLM Stats Score (conservative rating)80.6
SEAL - MultiNRC49Score77.9
SEAL - MASK86.33Score77.3
SEAL - TutorBench54.09Score76.9
SEAL - Professional Reasoning Benchmark - Legal49.33Score74.2
EnigmaEval11.23Score (self-reported)74
SEAL - Humanity's Last Exam23.68Score74
SEAL - Professional Reasoning Benchmark - Finance48.01Score71

Interactive version: theaggregate.ai/model?slug=gpt-5-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.