Context Arena MRCR (2-needle) — leaderboard

OpenAI MRCR benchmark: finds and distinguishes between 2 identical pieces of information in long conversations (8K–1M tokens). Tests multi-round co-reference resolution via AUC across context lengths.

Metric: AUC@1M (%). Source: contextarena.ai. Status: saturated. 89 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview) (High)93.8
2Gemini 3 Flash (Preview) (Medium)90
3Gemini 3.1 Pro (Preview) (High)87.8
4Gemini 3.1 Pro (Preview) (Low)87.4
5Gemini 3 Pro (Preview)85.5
6Gemini 3 Flash (Preview)85.5
7Gemini 2.5 Flash (Preview 05-20)78.3
8Gemini 2.5 Pro (Preview 06-05)77.5
9nova-2-lite-v163.5
10GPT-4.153.2
11GPT-4.1 Mini43.6
12Grok 4.1 Fast40.3
13Llama 4 Maverick39.9
14Gemini 2.5 Flash Lite (Preview 09-2025)34.3
15Gemini 2.0 Flash (001)32.1

Interactive version: theaggregate.ai/benchmark?slug=context-arena-mrcr-2-needle · How the rankings work · Data refreshed daily, snapshot 2026-07-22.