MiroEval (Multimodal): leaderboard
Metric: Overall (0-100), the mean of synthesis, factuality and process scores, on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | MiroThinker-H1 | 74.5 | |
| 2 | MiroThinker-1.7 | 71.6 | |
| 3 | OpenAI Deep Research | 70.2 | |
| 4 | Gemini 3.1 Pro Deep Research | 68.1 | |
| 5 | Claude Research (Opus 4.6) | 66.4 | |
| 6 | ChatGLM Agent | 63.6 | |
| 7 | MiniMax-M2.5 Research | 63.3 | |
| 8 | Manus-1.6-Max Wide Research | 62 | |
| 9 | Grok Deep Research | 60.5 | |
| 10 | Qwen-3.5-Plus Deep Research | 56.1 |
Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal · How It Works · Data refreshed daily, snapshot 2026-10-11.