GPT-5.6 Luna (xHigh): benchmark results

GPT-5.6 Luna evaluated at the xhigh reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1692 ± 1, rank #125 of 1761 rated models, from 76 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Chatbot Arena (Text - German)1486Arena Score94.2
AA Long Context Reasoning81.67Accuracy (%)94
Chatbot Arena (Text - Math)1478Arena Score93.5
LLM Chess (Saplin)1185.3ELO93.2
AA CritPt20.57Accuracy (%)93
Artificial Analysis Intelligence Index41.58Intelligence Index91.2
Chess Bench LLM1469Lichess Rating90.6
Chatbot Arena (Text - Japanese)1449Arena Score89.4
AA Humanity's Last Exam36.98Accuracy (%)87.9
AA Omniscience - Software Engineering (SWE)65.9Accuracy (%)87.7
AA GPQA Diamond89.49Accuracy (%)87.6
AA Omniscience - Science, Engineering & Mathematics45Accuracy (%)86.7

Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-xhigh · How It Works · Data refreshed daily, snapshot 2026-09-05.