GPT-5.2 (Thinking): benchmark results
Provider: OpenAI. Released 2025-12-11. Access: API.
Unified ELO 1697 ± 1, rank #146 of 3078 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| InteractBench - Medium - pass@1 | 53 | pass@1 (%; n=10 samples per task) | 100 |
| OpenAI GPT-5.6 System Card - Challenging Prompts - Gore | 87.7 | not_unsafe rate (%) | 100 |
| OpenAI GPT-5.6 System Card - Challenging Prompts - Violent Illicit Behavior | 97.5 | not_unsafe rate (%) | 100 |
| TRIP-Bench - Overall - Loose | 45 | Loose Success (%) | 100 |
| TRIP-Bench - Overall - Strict | 18.5 | Strict Success (%) | 100 |
| K-MetBench | 87.8 | Accuracy (self-reported) | 98.3 |
| InteractBench - Easy - pass@1 | 74.2 | pass@1 (%; n=10 samples per task) | 93.3 |
| InteractBench - Easy - pass@5 | 92.1 | pass@5 (%; n=10 samples per task) | 93.3 |
| InteractBench - Hard - pass@1 | 34 | pass@1 (%; n=10 samples per task) | 93.3 |
| InteractBench - Hard - pass@5 | 49.2 | pass@5 (%; n=10 samples per task) | 93.3 |
| InteractBench - Medium - pass@5 | 69 | pass@5 (%; n=10 samples per task) | 93.3 |
| InteractBench - Overall - pass@1 | 54.2 | pass@1 (%; n=10 samples per task) | 93.3 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.