GPT-5.1 Codex Mini (High) — benchmark results
Provider: OpenAI. Released 2025-11-19. Access: API.
Unified ELO 1729 ± 23, rank #226 of 1841 rated models, from 53 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA LiveCodeBench | 83.6 | Pass@1 (%) | 94.7 |
| AA AIME 2025 | 91.67 | Accuracy (%) | 93.1 |
| LLM Chess (Saplin) | 544 | ELO | 84.3 |
| AA IFBench | 67.89 | Accuracy (%) | 81.4 |
| Epoch AI - Scicode | 42.59 | Score | 78.9 |
| AA MMLU-Pro | 82 | Accuracy (%) | 77.9 |
| AA Terminal-Bench Hard | 33.33 | Accuracy (%) | 77.7 |
| Artificial Analysis Intelligence Index | 30.63 | Intelligence Index | 75.8 |
| AA Long Context Reasoning | 62.67 | Accuracy (%) | 75.2 |
| AA Global-MMLU-Lite - Yoruba | 69.92 | Accuracy (%) | 74.8 |
| AA GPQA Diamond | 81.31 | Accuracy (%) | 74.7 |
| AA Humanity's Last Exam | 16.91 | Accuracy (%) | 74.7 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-codex-mini-high · How It Works · Data refreshed daily, snapshot 2026-07-25.