GPT-5.1 Codex (High) — benchmark results

GPT-5.1 Codex evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-11-19. Access: API.

Unified ELO 1830 ± 27, rank #114 of 1776 rated models, from 59 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA AIME 202595.67Accuracy (%)97.2
AA LiveCodeBench84.87Pass@1 (%)95.9
AA MMLU-Pro86.01Accuracy (%)94.5
AA Global-MMLU-Lite - French92.75Accuracy (%)94
AA Global-MMLU-Lite - Hindi90.58Accuracy (%)94
AA Global-MMLU-Lite - Swahili89.5Accuracy (%)93.2
AA Global-MMLU-Lite - Japanese91.08Accuracy (%)92.2
AA Omniscience - Health39Accuracy (%)92.2
AA Omniscience - Software Engineering (SWE) - Swift72Accuracy (%)91.7
AA Global-MMLU-Lite - Spanish92.75Accuracy (%)90.6
AA Global-MMLU-Lite - Bengali89.92Accuracy (%)89.8
ALE-Bench1244.92Performance (Self-Refine x1) (self-reported)89.7

Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.