GPT-5.5 (Medium): benchmark results
GPT-5.5 evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2026-04-23. Access: API.
Unified ELO 1718 ± 1, rank #67 of 1761 rated models, from 46 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI - Natural Intelligence | 77.12 | NatInt Score | 99.8 |
| LLM Chess (Saplin) | 1532.2 | ELO | 98.8 |
| AA Omniscience - Software Engineering (SWE) | 85.5 | Accuracy (%) | 97.4 |
| AA Terminal-Bench Hard | 57.58 | Accuracy (%) | 97.3 |
| AA Long Context Reasoning | 83 | Accuracy (%) | 97 |
| HalluHard | 0.42 | Turn-1 Hallucination Rate | 96.8 |
| LisanBench | 0.64 | Mean Path Length / Current Maximum | 96.1 |
| AA Omniscience - Health | 49.75 | Accuracy (%) | 95.9 |
| AA Omniscience - Business | 47.6 | Accuracy (%) | 95.7 |
| AA Omniscience - Humanities & Social Sciences | 54.3 | Accuracy (%) | 95.7 |
| AA-Omniscience Accuracy | 56.8 | Accuracy (%) | 95.7 |
| UGI - Writing | 64.75 | Writing Score | 95.4 |
Interactive version: theaggregate.ai/model?slug=gpt-5-5-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.