BABILong (NIAH, v0, 0k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 30 models tracked.

Top models

#ModelScore
1GPT-4 Preview (0125)87
2Llama 3.1 70B Instruct85
3Mixtral 8x22B Instruct (v0.1)75
4Yi 34B (Base)72
5Phi-3-medium-128k-instruct72
6mamba-2.8B-hf70
7Llama 3.1 8B Instruct67
8Mixtral 8x7B Instruct (v0.1)65
9Yi-34B-200K65
10Llama 3 8B Instruct64
11Phi-3 Mini 128K Instruct64
12c4ai-command-r-v0164
13Mistral 7B Instruct (v0.2)60
14RWKV-6-World-7B56
15Yi-9B-200K52

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-0k · How It Works · Data refreshed daily, snapshot 2026-10-09.