GPT-5.6 Sol: benchmark results
OpenAI's flagship GPT-5.6 tier for demanding reasoning, coding, and agentic tasks. Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1761 ± 1, rank #14 of 1392 rated models, from 298 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI for Education Visual Maths | 91.14 | Accuracy (%) | 100 |
| AI for Education Visual Maths - Number and Operations | 89.19 | Accuracy (%) | 100 |
| AI for Education Visual Reasoning - odd one out | 88.3 | Accuracy (%) | 100 |
| Agentic Commerce World | 85.6 | Overall (%) | 100 |
| AgenticVBench | 38.4 | Average Success (%) | 100 |
| BenchX - Pass Rate | 82.94 | Pass Rate (%) | 100 |
| BoundaryBench (Unrestricted) | 83.9 | Success Rate (%) | 100 |
| Creative Writing v3 | 2208 | Elo score (self-reported) | 100 |
| DecBench | 57.2 | Exact Recovery (%) | 100 |
| ExploitGym | 293 | Successful Intended Exploits (#) | 100 |
| Graphwalks BFS 1M F1 | 83.4 | F1 (self-reported) | 100 |
| Graphwalks BFS 256k F1 | 95.4 | F1 (self-reported) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-sol · How It Works · Data refreshed daily, snapshot 2026-09-05.