GPT-5.6 Pro Luna: benchmark results

Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1947 ± 29, rank #37 of 2656 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Chess Leaderboard (Reasoning)1697Elo96.9
Tinybird AI SQL Benchmark - First-Attempt Success Rate100Questions answered with a valid query on the first attempt (93.4
LM Market Cap LMC Score89LMC Score (0-100)92
OpenRouter GPQA Diamond90.3Accuracy (%)89.1
Tinybird AI SQL Benchmark - Exactness53.47Result exactness vs human reference queries (0-100)79.1
MineBench1778Elo Rating74.2
Tinybird AI SQL Benchmark - Success Rate100Questions answered with a valid query within 3 retries (%)72.5
OpenRouter Tau2-Bench Airline71.3Accuracy (%)56.2
PRISM (1C:Enterprise) - Algorithmic Tasks (A)73Algorithmic BSL tasks fully solved, run in OneScript (%)54.9
PRISM (1C:Enterprise) - Platform Tasks (B)651C platform tasks fully solved, run in headless 1C (%)47.1

Interactive version: theaggregate.ai/model?slug=gpt-5-6-pro-luna · How It Works · Data refreshed daily, snapshot 2026-09-19.