BABILong (NIAH, v0, 1k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 30 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct81
2GPT-4 Preview (0125)81
3Mixtral 8x22B Instruct (v0.1)73
4Phi-3-medium-128k-instruct70
5Llama 3.1 8B Instruct68
6c4ai-command-r-v0164
7Mixtral 8x7B Instruct (v0.1)63
8Llama 3 8B Instruct60
9Yi-34B-200K59
10Phi-3 Mini 128K Instruct57
11Mistral 7B Instruct (v0.2)56
12Yi-9B-200K55
13RWKV-6-World-7B55
14Llama 2 7B 32K53
15Yi 34B (Base)52

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-1k · How It Works · Data refreshed daily, snapshot 2026-10-09.