NoLiMa — leaderboard

Long-context benchmark that goes beyond literal needle matching. Tests whether models can truly understand and retrieve information from long documents, not just pattern-match on exact strings.

Metric: Base Score (%). Source: github.com. Status: saturation imminent. 22 models tracked.

Top models

#ModelScore
1GPT-4o99.3
2Llama 3.3 70B97.3
3GPT-4.197
4Llama 3.1 405B94.7
5Llama 3.1 70B94.5
6Gemini 1.5 Pro92.6
7Jamba 1.5 Mini92.4
8Command-R+90.9
9Llama 4 Maverick90.1
10Gemini 2.0 Flash89.4
11Gemma 3 27B88.6
12Mistral Large 2 (Jul)87.9
13Claude 3.5 Sonnet87.6
14Gemma 3 12B87.4
15GPT-4o Mini84.9

Interactive version: theaggregate.ai/benchmark?slug=nolima · How the rankings work · Data refreshed daily, snapshot 2026-07-22.