BABILong (NIAH, v0, 8k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 29 models tracked.

Top models

#ModelScore
1GPT-4 Preview (0125)71
2Llama 3.1 70B Instruct70
3Llama 3.1 8B Instruct62
4Phi-3-medium-128k-instruct60
5c4ai-command-r-v0159
6Mixtral 8x22B Instruct (v0.1)58
7Yi-34B-200K52
8Mixtral 8x7B Instruct (v0.1)50
9Phi-3 Mini 128K Instruct50
10Mistral 7B Instruct (v0.2)45
11Yi-9B-200K45
12Llama 3 8B Instruct44
13LongChat-7B-v1.5-32k42
14Llama 2 7B 32K Instruct40
15Llama 2 7B 32K39

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-8k · How It Works · Data refreshed daily, snapshot 2026-10-09.