GPT-5.1 (2025-11-13): benchmark results
Provider: OpenAI. Released 2025-11-12. Access: API.
Unified ELO 1765 ± 10, rank #95 of 1639 rated models, from 777 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LiveFact - Inference Macro-F1 | 81.99 | Macro-F1 (%) in Inference Mode (label adjusted to what the e | 100 |
| RuleWorld - Single-Rule (Natural Language) | 100 | Exact-match accuracy (%; x100 of the 0-1 score; single-rule | 100 |
| Nejumi 4 - jaster (0-shot) - JCoLA (in-domain) | 84 | Exact match (%) | 99.3 |
| Horangi 4 - HRM8K | 97 | Accuracy (%) | 99 |
| Horangi 4 - IFEval-Ko | 93.75 | Instruction-following accuracy (%) | 99 |
| MedQA | 96.38 | Score (%) | 98.9 |
| Vals AI MedQA | 96.38 | Accuracy (%) | 98.9 |
| HELM AIR-Bench 2024 - #39-40.25: Characterization of identity - Sexual orientation | 80 | Refusal Rate (%) | 98.8 |
| HELM AIR-Bench 2024 - #39-40.26: Characterization of identity - Religion | 100 | Refusal Rate (%) | 98.8 |
| Nejumi 4 - jaster (2-shot) - AIO | 96.18 | Character F1 (x100) | 98.6 |
| HELM Arabic Enterprise - Arabic Legal Rag | 99 | Score | 98.5 |
| HELM Capabilities - WildBench | 86.29 | WB Score | 98.5 |
Interactive version: theaggregate.ai/model?slug=gpt-5-1-2025-11-13 · How It Works · Data refreshed daily, snapshot 2026-10-09.