GPT-5.1 (2025-11-13): benchmark results

Provider: OpenAI. Released 2025-11-12. Access: API.

Unified ELO 1765 ± 10, rank #95 of 1639 rated models, from 777 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LiveFact - Inference Macro-F181.99Macro-F1 (%) in Inference Mode (label adjusted to what the e100
RuleWorld - Single-Rule (Natural Language)100Exact-match accuracy (%; x100 of the 0-1 score; single-rule 100
Nejumi 4 - jaster (0-shot) - JCoLA (in-domain)84Exact match (%)99.3
Horangi 4 - HRM8K97Accuracy (%)99
Horangi 4 - IFEval-Ko93.75Instruction-following accuracy (%)99
MedQA96.38Score (%)98.9
Vals AI MedQA96.38Accuracy (%)98.9
HELM AIR-Bench 2024 - #39-40.25: Characterization of identity - Sexual orientation80Refusal Rate (%)98.8
HELM AIR-Bench 2024 - #39-40.26: Characterization of identity - Religion100Refusal Rate (%)98.8
Nejumi 4 - jaster (2-shot) - AIO96.18Character F1 (x100)98.6
HELM Arabic Enterprise - Arabic Legal Rag99Score98.5
HELM Capabilities - WildBench86.29WB Score98.5

Interactive version: theaggregate.ai/model?slug=gpt-5-1-2025-11-13 · How It Works · Data refreshed daily, snapshot 2026-10-09.