GPT-5.2 (Thinking): benchmark results

Provider: OpenAI. Released 2025-12-11. Access: API.

Unified ELO 1697 ± 1, rank #146 of 3078 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
InteractBench - Medium - pass@153pass@1 (%; n=10 samples per task)100
OpenAI GPT-5.6 System Card - Challenging Prompts - Gore87.7not_unsafe rate (%)100
OpenAI GPT-5.6 System Card - Challenging Prompts - Violent Illicit Behavior97.5not_unsafe rate (%)100
TRIP-Bench - Overall - Loose45Loose Success (%)100
TRIP-Bench - Overall - Strict18.5Strict Success (%)100
K-MetBench87.8Accuracy (self-reported)98.3
InteractBench - Easy - pass@174.2pass@1 (%; n=10 samples per task)93.3
InteractBench - Easy - pass@592.1pass@5 (%; n=10 samples per task)93.3
InteractBench - Hard - pass@134pass@1 (%; n=10 samples per task)93.3
InteractBench - Hard - pass@549.2pass@5 (%; n=10 samples per task)93.3
InteractBench - Medium - pass@569pass@5 (%; n=10 samples per task)93.3
InteractBench - Overall - pass@154.2pass@1 (%; n=10 samples per task)93.3

Interactive version: theaggregate.ai/model?slug=gpt-5-2-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.