GPT-6 Luna: benchmark results

Provider: OpenAI. Access: API.

Unified ELO 1935 ± 17, rank #50 of 2928 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SvelteBench100Average pass@1 (%)94.6
AI Chess Leaderboard (Continuation)1346Elo91.1
LLM Stats Score44.52LLM Stats Score (conservative rating)89.3
AI Chess Leaderboard (Reasoning)1067Elo79.9
BenchLM62.3Overall Score75.4
LLM Stats (Agents' Last Exam)50.9Score (%)70
RuneBench3741Total Peak XP Rate (XP/min)64.9
Featherbench96Pass Rate (%)62.5
ParseBench62.38Overall Score58.3
LLM Stats (DeepSWE 1.1)66.6Score (%)54.1
LLM Stats (FrontierCode 1.1)42.4Score (%)44.7
LLM Stats (OSWorld 2.0)52.7Score (%)38.5

Interactive version: theaggregate.ai/model?slug=gpt-6-luna · How It Works · Data refreshed daily, snapshot 2026-09-23.