MiroEval: leaderboard

Metric: Overall (0-100): 0.7 times the text-only overall plus 0.3 times the multimodal overall (each the mean of synthesis, factuality and process scores) over 100 deep-research tasks; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1MiroThinker-H176.6
2OpenAI Deep Research74.8
3MiroThinker-1.774.3
4Gemini 3.1 Pro Deep Research69.3
5Claude Research (Opus 4.6)67.3
6MiniMax-M2.5 Research66.2
7ChatGLM Agent65.1
8Manus-1.6-Max Wide Research63.4
9Qwen-3.5-Plus Deep Research62.1
10Grok Deep Research60.3

Interactive version: theaggregate.ai/benchmark?slug=miroeval · How It Works · Data refreshed daily, snapshot 2026-10-11.