GPT-5.6 Luna (Medium): benchmark results

GPT-5.6 Luna evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2026-07-09. Access: API.

Unified ELO 1659 ± 1, rank #214 of 1761 rated models, from 39 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Omniscience - Software Engineering (SWE)65.1Accuracy (%)87.3
AA Omniscience - Science, Engineering & Mathematics44.6Accuracy (%)85.8
AA-Omniscience Accuracy40.7Accuracy (%)83.8
AA Omniscience - Humanities & Social Sciences37.3Accuracy (%)82.5
AA Omniscience - Business32.1Accuracy (%)81.9
AA Omniscience - Health36.2Accuracy (%)81.9
Wolfram LLM Benchmarking Project56Correct Functionality (%)81.6
Artificial Analysis Intelligence Index30.19Intelligence Index81.2
AA GPQA Diamond85.86Accuracy (%)80.1
AA Omniscience - Law28.9Accuracy (%)79.9
LisanBench0.14Mean Path Length / Current Maximum79.7
LLM Chess (Saplin)730.1ELO79.5

Interactive version: theaggregate.ai/model?slug=gpt-5-6-luna-medium · How It Works · Data refreshed daily, snapshot 2026-09-05.