GPT-5.1 Codex Max — benchmark results
OpenAI's higher-compute GPT-5.1 Codex variant for coding and software-agent tasks. Provider: OpenAI. Released 2025-11-19. Access: API.
Unified ELO 1785 ± 28, rank #157 of 1776 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SnakeBench | 36.4 | TrueSkill Rating | 98.6 |
| AI Chess Leaderboard (Continuation) | 1354 | Elo | 92.6 |
| AI Chess Leaderboard (Reasoning) | 1354 | Elo | 90.6 |
| SWE-rebench | 54.48 | Resolved (%) | 80 |
| Terminal-Bench 2.0 | 60.4 | Accuracy (%) | 74.1 |
| Vals AI LiveCodeBench | 83.56 | Accuracy (%) | 72.4 |
| METR Benchmark (80% Horizon) | 0.84 | 80% Time Horizon (hours) | 72 |
| Vals AI IOI | 21.42 | Accuracy (%) | 69 |
| METR Benchmark | 3.73 | 50% Time Horizon (hours) | 68 |
| PostTrainBench | 19.7 | Weighted Avg Score | 63.6 |
| SRE Skills Bench - GMCQ | 89 | Accuracy (%) | 61.9 |
| SRE Skills Bench - Storage | 95 | Accuracy (%) | 61.9 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.