GPT-6 Luna: benchmark results
Provider: OpenAI. Access: API.
Unified ELO 1935 ± 17, rank #50 of 2928 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SvelteBench | 100 | Average pass@1 (%) | 94.6 |
| AI Chess Leaderboard (Continuation) | 1346 | Elo | 91.1 |
| LLM Stats Score | 44.52 | LLM Stats Score (conservative rating) | 89.3 |
| AI Chess Leaderboard (Reasoning) | 1067 | Elo | 79.9 |
| BenchLM | 62.3 | Overall Score | 75.4 |
| LLM Stats (Agents' Last Exam) | 50.9 | Score (%) | 70 |
| RuneBench | 3741 | Total Peak XP Rate (XP/min) | 64.9 |
| Featherbench | 96 | Pass Rate (%) | 62.5 |
| ParseBench | 62.38 | Overall Score | 58.3 |
| LLM Stats (DeepSWE 1.1) | 66.6 | Score (%) | 54.1 |
| LLM Stats (FrontierCode 1.1) | 42.4 | Score (%) | 44.7 |
| LLM Stats (OSWorld 2.0) | 52.7 | Score (%) | 38.5 |
Interactive version: theaggregate.ai/model?slug=gpt-6-luna · How It Works · Data refreshed daily, snapshot 2026-09-23.