MRCR v2 — leaderboard
Google DeepMind multi-round coreference benchmark that tests whether a model can disambiguate repeated requests and retrieve the intended item from long conversations.
Source: github.com. Status: saturation imminent.
Interactive version: theaggregate.ai/benchmark?slug=mrcr-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.