CorpusQA 1M — leaderboard

CorpusQA 1M is a long-context question answering benchmark designed to evaluate models at approximately 1 million token contexts. Models are scored on accuracy when retrieving and reasoning over information distributed across an extremely long input corpus.

Source: huggingface.co.

Interactive version: theaggregate.ai/benchmark?slug=corpusqa-1m · How the rankings work · Data refreshed daily, snapshot 2026-07-22.