MERRIN (Agentic Search): leaderboard
Metric: Accuracy (%) in the paper's smolagents multimodal search agent (Serper web search, plus page-reading and video-watching tools that call Gemini 3 Flash), on MERRIN's 162 human-annotated questions with unique short answers that need non-text web evidence (image, video, audio or table) and multi-hop reasoning or conflict resolution across noisy sources; accuracy against the gold answer by an LLM judge with the BrowseComp prompt, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 40.1 |
| 2 | Gemini 3 Pro | 39.9 |
| 3 | Gemini 3 Flash | 32.9 |
| 4 | GPT-5.4 Nano | 31.9 |
| 5 | GPT-5.4 Mini | 31.1 |
| 6 | Gemini 3.1 Flash Lite | 26.3 |
| 7 | Qwen 3 235B A22B | 23.3 |
| 8 | Qwen 3 30B A3B | 16.1 |
| 9 | Qwen 3 4B | 10.5 |
Interactive version: theaggregate.ai/benchmark?slug=merrin-agentic-search · How It Works · Data refreshed daily, snapshot 2026-10-07.