GPT-5.1 (Thinking) — benchmark results

GPT-5.1 evaluated with thinking enabled. Provider: OpenAI. Released 2025-11-12. Access: API.

Unified ELO 1783 ± 22, rank #159 of 1776 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SuperGPQA66.35Accuracy (%)97.1
ZeroEval GPQA Diamond88.1GPQA Diamond Score86.5
SEAL - EnigmaEval11.23Score83.1
SEAL - MultiChallenge63.41Score82.8
SEAL - Professional Reasoning Benchmark - Legal49.33Score82.1
SEAL - Professional Reasoning Benchmark - Finance48.01Score78.6
SEAL - MultiNRC49Score77.9
SEAL - MASK86.33Score77.3
SEAL - TutorBench54.09Score76.9
SEAL - Humanity's Last Exam23.68Score75.5
SEAL - Fortress25.72Score49.1
SEAL - VISTA43.82Score46.8

Interactive version: theaggregate.ai/model?slug=gpt-5-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.