GPT-5.1 Codex Mini (High): benchmark results

Provider: OpenAI. Released 2025-11-19. Access: API.

Unified ELO 1605 ± 1, rank #437 of 1919 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA IFBench67.89Accuracy (%)81.4
AA Terminal-Bench Hard33.33Accuracy (%)77.7
LLM Chess (Saplin)544ELO77
AA Global-MMLU-Lite - Yoruba69.92Accuracy (%)74.8
Artificial Analysis Intelligence Index20.38Intelligence Index71.4
AA Humanity's Last Exam18.49Accuracy (%)69.5
AA GPQA Diamond81.31Accuracy (%)69.4
AA Global-MMLU-Lite - Swahili80.08Accuracy (%)68.9
AA Omniscience - Health25.33Accuracy (%)67
AA Omniscience-16.42Score66.5
AA Global-MMLU-Lite85.02Accuracy (%)65.8
AA Global-MMLU-Lite - Arabic85.67Accuracy (%)65.8

Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex-mini-high · How It Works · Data refreshed daily, snapshot 2026-09-08.