MiroEval (Multimodal) - Synthesis: leaderboard
Metric: Synthesis quality (0-100): task-specific weighted rubric over coverage, insight, instruction following, clarity and query specification, scored by a GPT-series judge, on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | MiroThinker-H1 | 71.5 | |
| 2 | MiroThinker-1.7 | 69 | |
| 3 | OpenAI Deep Research | 66.7 | |
| 4 | Gemini 3.1 Pro Deep Research | 66.4 | |
| 5 | Claude Research (Opus 4.6) | 62.5 | |
| 6 | ChatGLM Agent | 61.6 | |
| 7 | MiniMax-M2.5 Research | 56.7 | |
| 8 | Grok Deep Research | 56.3 | |
| 9 | Manus-1.6-Max Wide Research | 54.3 | |
| 10 | Qwen-3.5-Plus Deep Research | 44.6 |
Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal-synthesis · How It Works · Data refreshed daily, snapshot 2026-10-11.