GPT-5.2 (xHigh) — benchmark results
GPT-5.2 evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2025-12-11. Access: API.
Unified ELO 2012 ± 26, rank #19 of 1776 rated models, from 109 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Context-Bench Skills | 85.31 | Task Completion (%) | 100 |
| Pencil Puzzle Bench - Firefly | 33.3 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - LITS | 53.3 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - Norinori | 93.3 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - Shikaku | 80 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - Sudoku | 20 | Direct-ask Success Rate (%) | 100 |
| AA AIME 2025 | 99 | Accuracy (%) | 99.6 |
| Pencil Puzzle Bench - Kurodoko | 6.7 | Direct-ask Success Rate (%) | 99 |
| Pencil Puzzle Bench - Mashu | 60 | Direct-ask Success Rate (%) | 99 |
| Pencil Puzzle Bench - Tapa | 60 | Direct-ask Success Rate (%) | 99 |
| MathArena - HMMT Feb 2025 | 100 | Accuracy (%) | 98.9 |
| AA LiveCodeBench | 88.89 | Pass@1 (%) | 98.5 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-xhigh · How the rankings work · Data refreshed daily, snapshot 2026-07-22.