MRCR v2: leaderboard

Google DeepMind multi-round coreference benchmark that tests whether a model can disambiguate repeated requests and retrieve the intended item from long conversations.

Source: github.com. Status: saturation imminent.

Interactive version: theaggregate.ai/benchmark?slug=mrcr-v2 · How It Works · Data refreshed daily, snapshot 2026-09-05.