GPT-5.2 (xHigh): benchmark results
GPT-5.2 evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2025-12-11. Access: API.
Unified ELO 1700 ± 1, rank #108 of 1761 rated models, from 110 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Context-Bench Skills | 85.31 | Task Completion (%) | 100 |
| Pencil Puzzle Bench - Firefly | 33.3 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - LITS | 53.3 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - Norinori | 93.3 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - Shikaku | 80 | Direct-ask Success Rate (%) | 100 |
| Pencil Puzzle Bench - Sudoku | 20 | Direct-ask Success Rate (%) | 100 |
| MathArena - HMMT Feb 2025 | 100 | Accuracy (%) | 99.3 |
| Pencil Puzzle Bench - Kurodoko | 6.7 | Direct-ask Success Rate (%) | 99 |
| Pencil Puzzle Bench - Mashu | 60 | Direct-ask Success Rate (%) | 99 |
| Pencil Puzzle Bench - Tapa | 60 | Direct-ask Success Rate (%) | 99 |
| MathArena - AIME 2025 | 100 | Accuracy (%) | 98.4 |
| Vals AI AIME | 96.88 | Accuracy (%) | 98.4 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.