GPT-5.2 (High) — benchmark results
GPT-5.2 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-12-11. Access: API.
Unified ELO 1978 ± 25, rank #33 of 1776 rated models, from 113 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CLEM AdventureGame | 99.17 | Game Clemscore (%) | 100 |
| CLEM Clean Up | 100 | Game Clemscore (%) | 100 |
| CLEM Codenames | 87.69 | Game Clemscore (%) | 100 |
| CLEM Deal or No Deal | 99.12 | Game Clemscore (%) | 100 |
| CLEM GuessWhat | 93.33 | Game Clemscore (%) | 100 |
| GDPval (OpenAI Evals) | 49.88 | Win Rate (%) | 100 |
| LLM2014 Logic 2025-12 | 81.83 | Median Score | 100 |
| LLM2014 Logic 2026-01 | 80.71 | Median Score | 100 |
| MCPMark | 57.48 | Pass@1 (%) | 100 |
| Pencil Puzzle Bench - Yajilin | 20 | Direct-ask Success Rate (%) | 100 |
| TriggerBench | 90.88 | Slot Match (Positive Clean) (self-reported) | 100 |
| Pencil Puzzle Bench - Slitherlink | 33.3 | Direct-ask Success Rate (%) | 99 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.