GPT-5.5 (Codex, xHigh): benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1833 ± 21, rank #65 of 1605 rated models, from 29 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ReFigBench (PPTX Skill Workflow) | 77.4 | Overall rubric score (0-100) under the specialized workflow: | 100 |
| ReFigBench (PPTX Skill Workflow) - Human Preference Elo | 1328.2 | Human preference Elo (specialized workflow: the agent drives | 100 |
| TUA-Bench | 64.7 | Success rate (%; mean of the verifier's scalar task reward o | 97.4 |
| Terminal-Bench 2.1 (Best Reported Harness) | 83.15 | Score (%) | 90.9 |
| SWE-Chain | 64.8 | F1-score (%; 2TP / (2TP + FP + FN) over each chain's release | 87.5 |
| GameXpert-Bench - GameFix - avg@3 (All Bugs Fixed) | 7.3 | Tasks with every injected bug fixed per run (of 100; mean of | 75 |
| GameXpert-Bench - GameFix - pass@3 (All Bugs Fixed) | 14 | Tasks with every injected bug fixed in at least one of three | 75 |
| ReFigBench (Direct Code Generation) | 74.2 | Overall rubric score (0-100) under the direct workflow: the | 75 |
| ReFigBench (Direct Code Generation) - Human Preference Elo | 1168.8 | Human preference Elo (direct workflow: the agent writes a pr | 75 |
| WebDev Arena (HTML) | 1545 | Arena Score | 74.2 |
| WebDev Arena (Simulations) | 1547 | Arena Score | 72.5 |
| WebDev Arena (Gaming) | 1558 | Arena Score | 71.4 |
Interactive version: theaggregate.ai/model?slug=gpt-5-5-codex-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-26.