GPT-5.1 Codex: benchmark results
OpenAI's Codex-specialized GPT-5.1 model for coding and software-agent workloads. Provider: OpenAI. Released 2025-11-19. Access: API.
Unified ELO 1668 ± 1, rank #101 of 1392 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PrepBench | 54.9 | End-to-End Prep-Code Acc. (self-reported) | 100 |
| SWE-Lancer | 66.3 | Score (%) | 100 |
| AI Chess Leaderboard (Reasoning) | 1743 | Elo | 97.5 |
| AI Chess Leaderboard (Continuation) | 1418 | Elo | 93.8 |
| SnakeBench | 32.1 | TrueSkill Rating | 90.9 |
| LM Market Cap LMC Score | 88.7 | LMC Score (0-100) | 88.5 |
| YapBench | 85.3 | YapIndex (lower is better) | 85.4 |
| MonitoringBench | 71.1 | Baseline catch rate (%) (self-reported) | 83.3 |
| BenchmarkList ECI | 134.35 | Capability Index (ECI) | 82.6 |
| BenchTable | 66.8 | Total Score (%) | 81.4 |
| WebApp1K Duo | 74.3 | Pass@1 (%) | 79.2 |
| Vals AI LiveCodeBench | 85.55 | Accuracy (%) | 78.9 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex · How It Works · Data refreshed daily, snapshot 2026-09-05.