MiroEval: leaderboard
Metric: Overall (0-100): 0.7 times the text-only overall plus 0.3 times the multimodal overall (each the mean of synthesis, factuality and process scores) over 100 deep-research tasks; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | MiroThinker-H1 | 76.6 | |
| 2 | OpenAI Deep Research | 74.8 | |
| 3 | MiroThinker-1.7 | 74.3 | |
| 4 | Gemini 3.1 Pro Deep Research | 69.3 | |
| 5 | Claude Research (Opus 4.6) | 67.3 | |
| 6 | MiniMax-M2.5 Research | 66.2 | |
| 7 | ChatGLM Agent | 65.1 | |
| 8 | Manus-1.6-Max Wide Research | 63.4 | |
| 9 | Qwen-3.5-Plus Deep Research | 62.1 | |
| 10 | Grok Deep Research | 60.3 |
Interactive version: theaggregate.ai/benchmark?slug=miroeval · How It Works · Data refreshed daily, snapshot 2026-10-11.