BABILong (NIAH, v0, 4k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 29 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct74
2GPT-4 Preview (0125)74
3Llama 3.1 8B Instruct66
4Mixtral 8x22B Instruct (v0.1)65
5Phi-3-medium-128k-instruct62
6c4ai-command-r-v0161
7Mixtral 8x7B Instruct (v0.1)55
8Yi-34B-200K54
9Phi-3 Mini 128K Instruct51
10Llama 3 8B Instruct50
11Mistral 7B Instruct (v0.2)49
12Yi-9B-200K46
13Llama 2 7B 32K Instruct43
14LongChat-7B-v1.5-32k41
15Llama 2 7B 32K40

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-4k · How It Works · Data refreshed daily, snapshot 2026-10-09.