BABILong (NIAH, v0, 2k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 29 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct78
2GPT-4 Preview (0125)77
3Mixtral 8x22B Instruct (v0.1)70
4Phi-3-medium-128k-instruct67
5Llama 3.1 8B Instruct66
6c4ai-command-r-v0163
7Mixtral 8x7B Instruct (v0.1)60
8Llama 3 8B Instruct58
9Yi-34B-200K56
10Phi-3 Mini 128K Instruct55
11Mistral 7B Instruct (v0.2)52
12Llama 2 7B 32K Instruct49
13Yi-9B-200K48
14RWKV-6-World-7B48
15Llama 2 7B 32K45

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-2k · How It Works · Data refreshed daily, snapshot 2026-10-09.