GPT-5.6 Pro Luna: benchmark results
Provider: OpenAI. Released 2026-07-09. Access: API.
Unified ELO 1947 ± 29, rank #37 of 2656 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Chess Leaderboard (Reasoning) | 1697 | Elo | 96.9 |
| Tinybird AI SQL Benchmark - First-Attempt Success Rate | 100 | Questions answered with a valid query on the first attempt ( | 93.4 |
| LM Market Cap LMC Score | 89 | LMC Score (0-100) | 92 |
| OpenRouter GPQA Diamond | 90.3 | Accuracy (%) | 89.1 |
| Tinybird AI SQL Benchmark - Exactness | 53.47 | Result exactness vs human reference queries (0-100) | 79.1 |
| MineBench | 1778 | Elo Rating | 74.2 |
| Tinybird AI SQL Benchmark - Success Rate | 100 | Questions answered with a valid query within 3 retries (%) | 72.5 |
| OpenRouter Tau2-Bench Airline | 71.3 | Accuracy (%) | 56.2 |
| PRISM (1C:Enterprise) - Algorithmic Tasks (A) | 73 | Algorithmic BSL tasks fully solved, run in OneScript (%) | 54.9 |
| PRISM (1C:Enterprise) - Platform Tasks (B) | 65 | 1C platform tasks fully solved, run in headless 1C (%) | 47.1 |
Interactive version: theaggregate.ai/model?slug=gpt-5-6-pro-luna · How It Works · Data refreshed daily, snapshot 2026-09-19.