MRCR v2 — leaderboard

Google DeepMind multi-round coreference benchmark that tests whether a model can disambiguate repeated requests and retrieve the intended item from long conversations.

Source: github.com. Status: saturation imminent.

Interactive version: theaggregate.ai/benchmark?slug=mrcr-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.