BABILong — leaderboard
BABILong evaluates long-context language models on synthetic QA tasks over increasingly long contexts, reporting accuracy by context length.
Source: huggingface.co. Status: saturation imminent.
Interactive version: theaggregate.ai/benchmark?slug=babilong · How the rankings work · Data refreshed daily, snapshot 2026-07-22.