GPT-5.5 (Codex, xHigh): benchmark results

Provider: OpenAI. Access: API.

Unified ELO 1833 ± 21, rank #65 of 1605 rated models, from 29 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ReFigBench (PPTX Skill Workflow)77.4Overall rubric score (0-100) under the specialized workflow:100
ReFigBench (PPTX Skill Workflow) - Human Preference Elo1328.2Human preference Elo (specialized workflow: the agent drives100
TUA-Bench64.7Success rate (%; mean of the verifier's scalar task reward o97.4
Terminal-Bench 2.1 (Best Reported Harness)83.15Score (%)90.9
SWE-Chain64.8F1-score (%; 2TP / (2TP + FP + FN) over each chain's release87.5
GameXpert-Bench - GameFix - avg@3 (All Bugs Fixed)7.3Tasks with every injected bug fixed per run (of 100; mean of75
GameXpert-Bench - GameFix - pass@3 (All Bugs Fixed)14Tasks with every injected bug fixed in at least one of three75
ReFigBench (Direct Code Generation)74.2Overall rubric score (0-100) under the direct workflow: the 75
ReFigBench (Direct Code Generation) - Human Preference Elo1168.8Human preference Elo (direct workflow: the agent writes a pr75
WebDev Arena (HTML)1545Arena Score74.2
WebDev Arena (Simulations)1547Arena Score72.5
WebDev Arena (Gaming)1558Arena Score71.4

Interactive version: theaggregate.ai/model?slug=gpt-5-5-codex-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-26.