GPT-5.5 (xHigh) — benchmark results
GPT-5.5 evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-04-23. Access: API.
Unified ELO 2057 ± 17, rank #13 of 1776 rated models, from 140 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - Rust | 92 | Accuracy (%) | 100 |
| ALE-Bench | 1942.97 | Performance (Self-Refine x1) (self-reported) | 100 |
| ClawProBench | 67.9 | Final Score (self-reported) | 100 |
| GRIPS | 96.4 | Accuracy (%) | 100 |
| GRIPS-hard | 87.3 | Accuracy (%) | 100 |
| LLM2014 Logic 2026-04 | 83.96 | Median Score | 100 |
| LLM2014 Logic 2026-05 | 80.47 | Median Score | 100 |
| LLM2014 Logic 2026-06 | 80.47 | Median Score | 100 |
| LLM2014 Logic 2026-07 | 77.46 | Median Score | 100 |
| LiveBench | 81.28 | LiveBench average (self-reported) | 100 |
| MathArena - APEX Shortlist 2025 | 98.4 | Accuracy (%) | 100 |
| MathArena - ARXIV March | 77.5 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-5-5-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.