GPT-5.2 (2025-12-11) (Non-reasoning): benchmark results
Provider: OpenAI. Released 2025-12-11. Access: API.
Unified ELO 1626 ± 21, rank #617 of 2131 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GPTNT - Manual VQA Element Grounding | 90 | Accuracy (%; 60 questions locating a described visual elemen | 100 |
| GR-Ben - Mixed-Form Reasoning | 56.1 | F1 (%) of first-error identification: the harmonic mean of t | 95.2 |
| GR-Ben - Biology | 48.9 | F1 (%) of first-error identification: the harmonic mean of t | 90.5 |
| GR-Ben - Computer Science | 53.5 | F1 (%) of first-error identification: the harmonic mean of t | 90.5 |
| UGI - Natural Intelligence | 50.84 | NatInt Score | 89 |
| CADEngBench - Functional Editing (L2-E) | 70.3 | Edit pass rate (%; share of 300 functional-edit tasks on sup | 85.7 |
| GR-Ben - Chemistry | 41.7 | F1 (%) of first-error identification: the harmonic mean of t | 85.7 |
| GR-Ben - Inductive Reasoning | 54.5 | F1 (%) of first-error identification: the harmonic mean of t | 85.7 |
| GR-Ben - Analogical Reasoning | 52.1 | F1 (%) of first-error identification: the harmonic mean of t | 83.3 |
| GR-Ben - Abductive Reasoning | 38.6 | F1 (%) of first-error identification: the harmonic mean of t | 81 |
| GR-Ben - Physics | 47.3 | F1 (%) of first-error identification: the harmonic mean of t | 81 |
| PLCC - Culture & Tradition | 86 | Accuracy (%) | 78.8 |
Interactive version: theaggregate.ai/model?slug=gpt-5-2-2025-12-11-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-09.