BABILong (NIAH, v0, 16k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct65
2GPT-4 Preview (0125)64
3Llama 3.1 8B Instruct60
4Phi-3-medium-128k-instruct57
5c4ai-command-r-v0152
6Mixtral 8x22B Instruct (v0.1)51
7Yi-34B-200K50
8Mixtral 8x7B Instruct (v0.1)46
9Phi-3 Mini 128K Instruct46
10Mistral 7B Instruct (v0.2)42
11LongChat-7B-v1.5-32k39
12Yi-9B-200K36
13Llama 2 7B 32K Instruct35
14Llama 2 7B 32K32
15Yi 34B (Base)31

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-16k · How It Works · Data refreshed daily, snapshot 2026-10-09.