BABILong (NIAH, v0, 32k): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Estimated already saturated. 25 models tracked.

Top models

#ModelScore
1Llama 3.1 70B Instruct59
2Llama 3.1 8B Instruct56
3Phi-3-medium-128k-instruct53
4GPT-4 Preview (0125)53
5c4ai-command-r-v0151
6Yi-34B-200K48
7Mixtral 8x22B Instruct (v0.1)43
8Phi-3 Mini 128K Instruct42
9Mixtral 8x7B Instruct (v0.1)40
10Mistral 7B Instruct (v0.2)37
11Yi-9B-200K37
12Yarn-Mistral-7B-128k16
13LongChat-7B-v1.5-32k5
14Llama 2 7B 32K Instruct5
15Yi 34B (Base)4

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-32k · How It Works · Data refreshed daily, snapshot 2026-10-09.