GPT-5.1 Codex Max — benchmark results

OpenAI's higher-compute GPT-5.1 Codex variant for coding and software-agent tasks. Provider: OpenAI. Released 2025-11-19. Access: API.

Unified ELO 1785 ± 28, rank #157 of 1776 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SnakeBench36.4TrueSkill Rating98.6
AI Chess Leaderboard (Continuation)1354Elo92.6
AI Chess Leaderboard (Reasoning)1354Elo90.6
SWE-rebench54.48Resolved (%)80
Terminal-Bench 2.060.4Accuracy (%)74.1
Vals AI LiveCodeBench83.56Accuracy (%)72.4
METR Benchmark (80% Horizon)0.8480% Time Horizon (hours)72
Vals AI IOI21.42Accuracy (%)69
METR Benchmark3.7350% Time Horizon (hours)68
PostTrainBench19.7Weighted Avg Score63.6
SRE Skills Bench - GMCQ89Accuracy (%)61.9
SRE Skills Bench - Storage95Accuracy (%)61.9

Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.