GPT-5.1 Codex: benchmark results

OpenAI's Codex-specialized GPT-5.1 model for coding and software-agent workloads. Provider: OpenAI. Released 2025-11-19. Access: API.

Unified ELO 1668 ± 1, rank #101 of 1392 rated models, from 42 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PrepBench54.9End-to-End Prep-Code Acc. (self-reported)100
SWE-Lancer66.3Score (%)100
AI Chess Leaderboard (Reasoning)1743Elo97.5
AI Chess Leaderboard (Continuation)1418Elo93.8
SnakeBench32.1TrueSkill Rating90.9
LM Market Cap LMC Score88.7LMC Score (0-100)88.5
YapBench85.3YapIndex (lower is better)85.4
MonitoringBench71.1Baseline catch rate (%) (self-reported)83.3
BenchmarkList ECI134.35Capability Index (ECI)82.6
BenchTable66.8Total Score (%)81.4
WebApp1K Duo74.3Pass@1 (%)79.2
Vals AI LiveCodeBench85.55Accuracy (%)78.9

Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex · How It Works · Data refreshed daily, snapshot 2026-09-05.