GPT-5.2 (High): benchmark results
GPT-5.2 evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-12-11. Access: API.
Unified ELO 1691 ± 1, rank #127 of 1761 rated models, from 156 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CLEM AdventureGame | 99.17 | Game Clemscore (%) | 100 |
| CLEM Clean Up | 100 | Game Clemscore (%) | 100 |
| CLEM Codenames | 87.69 | Game Clemscore (%) | 100 |
| CLEM Deal or No Deal | 99.12 | Game Clemscore (%) | 100 |
| CLEM GuessWhat | 93.33 | Game Clemscore (%) | 100 |
| GDPval (OpenAI Evals) | 49.88 | Win Rate (%) | 100 |
| LLM2014 Logic 2025-12 | 81.83 | Median Score | 100 |
| LLM2014 Logic 2026-01 | 80.71 | Median Score | 100 |
| LiveCodeBench Pro Hard | 15.94 | Pass@1 (%) | 100 |
| LiveCodeBench Pro Medium | 52.11 | Pass@1 (%) | 100 |
| MCPMark | 57.48 | Pass@1 (%) | 100 |
| Pencil Puzzle Bench - Yajilin | 20 | Direct-ask Success Rate (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-high · How It Works · Data refreshed daily, snapshot 2026-09-05.