GPT-5.2 (2025-12-11) (Non-reasoning): benchmark results

Provider: OpenAI. Released 2025-12-11. Access: API.

Unified ELO 1626 ± 21, rank #617 of 2131 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
GPTNT - Manual VQA Element Grounding90Accuracy (%; 60 questions locating a described visual elemen100
GR-Ben - Mixed-Form Reasoning56.1F1 (%) of first-error identification: the harmonic mean of t95.2
GR-Ben - Biology48.9F1 (%) of first-error identification: the harmonic mean of t90.5
GR-Ben - Computer Science53.5F1 (%) of first-error identification: the harmonic mean of t90.5
UGI - Natural Intelligence50.84NatInt Score89
CADEngBench - Functional Editing (L2-E)70.3Edit pass rate (%; share of 300 functional-edit tasks on sup85.7
GR-Ben - Chemistry41.7F1 (%) of first-error identification: the harmonic mean of t85.7
GR-Ben - Inductive Reasoning54.5F1 (%) of first-error identification: the harmonic mean of t85.7
GR-Ben - Analogical Reasoning52.1F1 (%) of first-error identification: the harmonic mean of t83.3
GR-Ben - Abductive Reasoning38.6F1 (%) of first-error identification: the harmonic mean of t81
GR-Ben - Physics47.3F1 (%) of first-error identification: the harmonic mean of t81
PLCC - Culture & Tradition86Accuracy (%)78.8

Interactive version: theaggregate.ai/model?slug=gpt-5-2-2025-12-11-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-09.