BABILong (NIAH, v0, 10M): leaderboard

Metric: QA1–5 Mean Accuracy (%). Source: huggingface.co. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1~ ARMT (137M) fine-tune76.6
2Llama3-ChatQA-1.5-8B + RAG37
3~ RMT (137M) fine-tune33.78

Interactive version: theaggregate.ai/benchmark?slug=babilong-niah-v0-10m · How It Works · Data refreshed daily, snapshot 2026-10-09.