LUNAR - Full Context: leaderboard

Metric: Personalization score (1-5; mean of coverage and depth, GPT-5.1 judge). Source: arxiv.org. 19 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)3.9
2Kimi K2.63.83
3DeepSeek V4 Flash3.57
4GLM-5.13.44
5MiniMax-M2.73.32
6DeepSeek V4 Pro3.24
7GPT-4.1 Mini3.11
8Ling-2.6-1T3.06
9Qwen 3 30B A3B 2507 Instruct2.91
10GLM-4.7 Flash2.85
11GPT-4o Mini2.76
12Ling-2.6-flash2.76
13Qwen 3 14B (Non-reasoning)2.6
14Qwen 3 Next 80B A3B Instruct2.58
15Qwen 3 8B (Non-reasoning)2.56

Interactive version: theaggregate.ai/benchmark?slug=lunar-full-context · How It Works · Data refreshed daily, snapshot 2026-09-19.