GPT-5.1 (High) — benchmark results
GPT-5.1 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-11-12. Access: API.
Unified ELO 1806 ± 14, rank #137 of 1776 rated models, from 123 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Long Context Reasoning | 75 | Accuracy (%) | 99.6 |
| AA Global-MMLU-Lite - French | 93.5 | Accuracy (%) | 99.1 |
| Vals AI CaseLaw v2 | 73.42 | Accuracy (%) | 98.3 |
| MedScribe | 88.09 | Score (self-reported) | 98.1 |
| AA LiveCodeBench | 86.77 | Pass@1 (%) | 97.5 |
| CLBench | 23.7 | Solving Rate (%) | 97.1 |
| AA Global-MMLU-Lite - English | 94.08 | Accuracy (%) | 96.6 |
| AA MMLU-Pro | 87.04 | Accuracy (%) | 96.5 |
| LLM2014 Logic 2025-11 | 79.53 | Median Score | 96.2 |
| LLM2014 Logic 2025-12 | 77.54 | Median Score | 96 |
| AA Global-MMLU-Lite - Bengali | 90.92 | Accuracy (%) | 95.8 |
| AA Global-MMLU-Lite - Swahili | 89.58 | Accuracy (%) | 95 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.