MiGUE-Bench - Event Prediction: leaderboard
Metric: Accuracy (%) on 300 next-event prediction multiple-choice questions with LLM-built distractors over news document sequences; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 58.14 |
| 2 | Claude Opus 4.5 | 55.44 |
| 3 | DeepSeek V3.2 | 52.1 |
| 4 | Kimi K2 | 51.5 |
| 5 | GLM-4.7 | 48.67 |
| 6 | Qwen 3 Max | 45.85 |
| 7 | Claude Haiku 4.5 | 44.52 |
| 8 | Qwen 3 235B A22B | 43.52 |
| 9 | GPT-5.2 Pro | 42.31 |
| 10 | Qwen 3 30B A3B | 35.22 |
| 11 | Qwen 3 8B | 33.55 |
Interactive version: theaggregate.ai/benchmark?slug=migue-bench-event-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.