GPT-5.6 Sol — benchmark results
OpenAI's flagship GPT-5.6 tier for demanding reasoning, coding, and agentic tasks. Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1978 ± 22, rank #32 of 1776 rated models, from 129 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI for Education Visual Reasoning - match (process) | 88.9 | Accuracy (%) | 100 |
| AI for Education Visual Reasoning - odd one out | 88.3 | Accuracy (%) | 100 |
| AI for Education Visual Reasoning - pattern completion (linear) | 92.3 | Accuracy (%) | 100 |
| AgenticVBench | 38.4 | Average Success (%) | 100 |
| Agents' Last Exam | 30.6 | Pass Rate (%) | 100 |
| Benchmarks.bio - SpatialBench-Long | 38.89 | Pass Rate (%) | 100 |
| Benchmarks.bio - scBench-Long | 38.1 | Pass Rate (%) | 100 |
| DeepSWE | 72.7 | Pass@1 (%) | 100 |
| ExploitGym | 293 | Successful Intended Exploits (#) | 100 |
| LLM Stats (Artificial Analysis) | 59 | Score (%) | 100 |
| LLM Stats (DeepSWE) | 72.7 | Score (%) | 100 |
| LLM Stats (Graphwalks BFS >128k) | 90.7 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-sol · How the rankings work · Data refreshed daily, snapshot 2026-07-22.