MiroEval (Multimodal) - Synthesis: leaderboard

Metric: Synthesis quality (0-100): task-specific weighted rubric over coverage, insight, instruction following, clarity and query specification, scored by a GPT-series judge, on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1MiroThinker-H171.5
2MiroThinker-1.769
3OpenAI Deep Research66.7
4Gemini 3.1 Pro Deep Research66.4
5Claude Research (Opus 4.6)62.5
6ChatGLM Agent61.6
7MiniMax-M2.5 Research56.7
8Grok Deep Research56.3
9Manus-1.6-Max Wide Research54.3
10Qwen-3.5-Plus Deep Research44.6

Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal-synthesis · How It Works · Data refreshed daily, snapshot 2026-10-11.