MiroEval (Multimodal) - Process: leaderboard

Metric: Process quality (0-100): audit of the research trajectory (search breadth, analytical depth, progressive refinement, critical thinking, efficiency) on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1MiroThinker-H173.5
2MiroThinker-1.767.4
3OpenAI Deep Research66.8
4Claude Research (Opus 4.6)65.9
5Gemini 3.1 Pro Deep Research64.1
6MiniMax-M2.5 Research62.2
7Manus-1.6-Max Wide Research61.8
8ChatGLM Agent57.7
9Grok Deep Research53.9
10Qwen-3.5-Plus Deep Research53.8

Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal-process · How It Works · Data refreshed daily, snapshot 2026-10-11.