MRCR-v2 128k: leaderboard

MRCR-v2 long-context retrieval subset using a 128K context window with 8 needles, as reported in Qwen's Qwen3.7-Max launch post.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 20 models tracked.

Top models

#ModelScore
1Gemini 3.7 Flash97
2GPT-5.6 Terra93.5
3Gemini 3.6 Flash91.8
4Qwen 3.7 Plus91.7
5Qwen 3.7 Max90.4
6Qwen 3.6 Plus85.9
7Gemini 3.1 Pro (Preview)84.9
8Claude Opus 4.6 (Max)84
9Claude Sonnet 581.5
10Grok 4.581.4
11Gemini 3.5 Flash77.3
12GPT-5.6 Luna74.8
13DeepSeek V4 Pro (Max)74.4
14Gemma 4 31B66.4
15Kimi K2.6 (Thinking)63.1

Interactive version: theaggregate.ai/benchmark?slug=mrcr-v2-128k · How It Works · Data refreshed daily, snapshot 2026-09-05.