GPT-5.1 Codex — benchmark results
OpenAI's Codex-specialized GPT-5.1 model for coding and software-agent workloads. Provider: OpenAI. Released 2025-11-19. Access: API.
Unified ELO 1755 ± 37, rank #188 of 1776 rated models, from 39 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PrepBench | 54.9 | End-to-End Prep-Code Acc. (self-reported) | 100 |
| SWE-Lancer | 66.3 | Score (%) | 100 |
| AI Chess Leaderboard (Reasoning) | 1743 | Elo | 98 |
| AI Chess Leaderboard (Continuation) | 1418 | Elo | 95.1 |
| CritPt | 5.7 | Accuracy (self-reported) | 91 |
| SnakeBench | 32.1 | TrueSkill Rating | 90.9 |
| AA-LCR | 67.3 | Score (self-reported) | 90.2 |
| Vals AI LiveCodeBench | 85.55 | Accuracy (%) | 84.3 |
| MonitoringBench | 71.1 | Baseline catch rate (%) (self-reported) | 83.3 |
| BenchTable | 66.8 | Total Score (%) | 81.4 |
| WebApp1K Duo | 74.3 | Pass@1 (%) | 79.2 |
| Terminal-Bench 2.0 | 57.8 | Accuracy (%) | 70.7 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex · How the rankings work · Data refreshed daily, snapshot 2026-07-22.