GPT-5.1 (Medium) — benchmark results

GPT-5.1 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-11-12. Access: API.

Unified ELO 1820 ± 19, rank #126 of 1776 rated models, from 107 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - CareQA Open89.02Score (%)100
Medmarks - MEDEC70.46Score (%)100
Medmarks - SCTpublic77.96Score (%)100
Medmarks - HEAD-QA v292.33Score (%)98.6
Medmarks - MedConceptsQA Easy99.97Score (%)98.6
Medmarks - MedHallu Hard60.85Score (%)98.6
Medmarks - MedHallu Medium74.06Score (%)98.6
Medmarks - MedMCQA84.65Score (%)98.6
Medmarks - MetaMedQA82.74Score (%)98.6
Medmarks - SuperGPQA Medicine Easy68.87Score (%)98.6
Medmarks - SuperGPQA Medicine Hard52.38Score (%)98.6
Medmarks - Medbullets OP587.23Score (%)97.9

Interactive version: theaggregate.ai/model?slug=gpt-5-1-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.