MRCR v2: leaderboard
Google DeepMind multi-round coreference benchmark that tests whether a model can disambiguate repeated requests and retrieve the intended item from long conversations.
Source: github.com. Status: saturation imminent.
Interactive version: theaggregate.ai/benchmark?slug=mrcr-v2 · How It Works · Data refreshed daily, snapshot 2026-09-05.