MiroEval (Multimodal) - Process: leaderboard
Metric: Process quality (0-100): audit of the research trajectory (search breadth, analytical depth, progressive refinement, critical thinking, efficiency) on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | MiroThinker-H1 | 73.5 | |
| 2 | MiroThinker-1.7 | 67.4 | |
| 3 | OpenAI Deep Research | 66.8 | |
| 4 | Claude Research (Opus 4.6) | 65.9 | |
| 5 | Gemini 3.1 Pro Deep Research | 64.1 | |
| 6 | MiniMax-M2.5 Research | 62.2 | |
| 7 | Manus-1.6-Max Wide Research | 61.8 | |
| 8 | ChatGLM Agent | 57.7 | |
| 9 | Grok Deep Research | 53.9 | |
| 10 | Qwen-3.5-Plus Deep Research | 53.8 |
Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal-process · How It Works · Data refreshed daily, snapshot 2026-10-11.