MiroEval (Multimodal): leaderboard

Metric: Overall (0-100), the mean of synthesis, factuality and process scores, on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1MiroThinker-H174.5
2MiroThinker-1.771.6
3OpenAI Deep Research70.2
4Gemini 3.1 Pro Deep Research68.1
5Claude Research (Opus 4.6)66.4
6ChatGLM Agent63.6
7MiniMax-M2.5 Research63.3
8Manus-1.6-Max Wide Research62
9Grok Deep Research60.5
10Qwen-3.5-Plus Deep Research56.1

Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal · How It Works · Data refreshed daily, snapshot 2026-10-11.