MiroEval (Multimodal) - Factuality: leaderboard

Metric: Factuality (0-100): share of the report's verifiable claims judged right by agentic verification against web sources and attachments, on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.

Top models

#ModelScoreOverall rank
1MiroThinker-H178.5
2MiroThinker-1.778.4
3OpenAI Deep Research77
4Gemini 3.1 Pro Deep Research73.7
5ChatGLM Agent71.6
6Grok Deep Research71.5
7MiniMax-M2.5 Research71
8Claude Research (Opus 4.6)70.7
9Manus-1.6-Max Wide Research70
10Qwen-3.5-Plus Deep Research69.9

Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal-factuality · How It Works · Data refreshed daily, snapshot 2026-10-11.