GPT-5.2 (xHigh) — benchmark results

GPT-5.2 evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2025-12-11. Access: API.

Unified ELO 2012 ± 26, rank #19 of 1776 rated models, from 109 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Context-Bench Skills85.31Task Completion (%)100
Pencil Puzzle Bench - Firefly33.3Direct-ask Success Rate (%)100
Pencil Puzzle Bench - LITS53.3Direct-ask Success Rate (%)100
Pencil Puzzle Bench - Norinori93.3Direct-ask Success Rate (%)100
Pencil Puzzle Bench - Shikaku80Direct-ask Success Rate (%)100
Pencil Puzzle Bench - Sudoku20Direct-ask Success Rate (%)100
AA AIME 202599Accuracy (%)99.6
Pencil Puzzle Bench - Kurodoko6.7Direct-ask Success Rate (%)99
Pencil Puzzle Bench - Mashu60Direct-ask Success Rate (%)99
Pencil Puzzle Bench - Tapa60Direct-ask Success Rate (%)99
MathArena - HMMT Feb 2025100Accuracy (%)98.9
AA LiveCodeBench88.89Pass@1 (%)98.5

Interactive version: theaggregate.ai/model?slug=gpt-5-2-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.