MiroEval (Multimodal) - Factuality: leaderboard
Metric: Factuality (0-100): share of the report's verifiable claims judged right by agentic verification against web sources and attachments, on the 30 multimodal deep-research tasks with image, PDF or spreadsheet attachments; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 10 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | MiroThinker-H1 | 78.5 | |
| 2 | MiroThinker-1.7 | 78.4 | |
| 3 | OpenAI Deep Research | 77 | |
| 4 | Gemini 3.1 Pro Deep Research | 73.7 | |
| 5 | ChatGLM Agent | 71.6 | |
| 6 | Grok Deep Research | 71.5 | |
| 7 | MiniMax-M2.5 Research | 71 | |
| 8 | Claude Research (Opus 4.6) | 70.7 | |
| 9 | Manus-1.6-Max Wide Research | 70 | |
| 10 | Qwen-3.5-Plus Deep Research | 69.9 |
Interactive version: theaggregate.ai/benchmark?slug=miroeval-multimodal-factuality · How It Works · Data refreshed daily, snapshot 2026-10-11.