GPT-5.1 (Medium) — benchmark results
GPT-5.1 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-11-12. Access: API.
Unified ELO 1820 ± 19, rank #126 of 1776 rated models, from 107 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - CareQA Open | 89.02 | Score (%) | 100 |
| Medmarks - MEDEC | 70.46 | Score (%) | 100 |
| Medmarks - SCTpublic | 77.96 | Score (%) | 100 |
| Medmarks - HEAD-QA v2 | 92.33 | Score (%) | 98.6 |
| Medmarks - MedConceptsQA Easy | 99.97 | Score (%) | 98.6 |
| Medmarks - MedHallu Hard | 60.85 | Score (%) | 98.6 |
| Medmarks - MedHallu Medium | 74.06 | Score (%) | 98.6 |
| Medmarks - MedMCQA | 84.65 | Score (%) | 98.6 |
| Medmarks - MetaMedQA | 82.74 | Score (%) | 98.6 |
| Medmarks - SuperGPQA Medicine Easy | 68.87 | Score (%) | 98.6 |
| Medmarks - SuperGPQA Medicine Hard | 52.38 | Score (%) | 98.6 |
| Medmarks - Medbullets OP5 | 87.23 | Score (%) | 97.9 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.