OpenAI GPT-6 Sol & Luna Launch - Factual Errors on User-Flagged Conversations: leaderboard

Metric: Answers with any factual error (%). Source: openai.com. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1GPT-6 (High)3.9
2GPT-6 (Max)3.91
3GPT-6 (xHigh)3.99
4GPT-6 (Medium)4.41
5GPT-6 Sol (xHigh)4.52
6GPT-6 Sol (Max)4.57
7GPT-6 Sol (High)5.14
8GPT-6 (Low)6.26
9GPT-6 Sol (Medium)6.86
10GPT-6 Luna (Max)7.56
11GPT-5.6 Sol (xHigh)8.44
12GPT-5.6 Sol (Max)8.5
13GPT-6 Luna (xHigh)10.15
14GPT-5.6 Sol (High)10.75
15GPT-6 Sol (Low)11.42

Interactive version: theaggregate.ai/benchmark?slug=openai-gpt-6-sol-luna-launch-factual-errors-on-user-flagged-conversations · How It Works · Data refreshed daily, snapshot 2026-09-24.