MiGUE-Bench - Intra-Document Coreference: leaderboard

Metric: Accuracy (%) on 300 single-document event coreference questions (2 options); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScore
1GPT-5.2 Pro97.35
2Claude Opus 4.597
3GLM-4.796.33
4DeepSeek V3.296.25
5Gemini 3 Pro95
6Claude Haiku 4.593
7Kimi K292
8Qwen 3 Max90.5
9Qwen 3 235B A22B87.25
10Qwen 3 8B86.77
11Qwen 3 30B A3B62.11

Interactive version: theaggregate.ai/benchmark?slug=migue-bench-intra-document-coreference · How It Works · Data refreshed daily, snapshot 2026-09-29.