MERRIN (No Search): leaderboard
Metric: Accuracy (%) with no search or browsing tool, on MERRIN's 162 human-annotated questions with unique short answers that need non-text web evidence (image, video, audio or table) and multi-hop reasoning or conflict resolution across noisy sources; accuracy against the gold answer by an LLM judge with the BrowseComp prompt, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 24.7 |
| 2 | Gemini 3 Pro | 23.5 |
| 3 | Gemini 3 Flash | 19.1 |
| 4 | GPT-5.4 Mini | 14 |
| 5 | Gemini 3.1 Flash Lite | 12.8 |
| 6 | Qwen 3 235B A22B | 12.1 |
| 7 | Qwen 3 4B | 10.3 |
| 8 | GPT-5.4 Nano | 9.9 |
| 9 | Qwen 3 30B A3B | 8 |
Interactive version: theaggregate.ai/benchmark?slug=merrin-no-search · How It Works · Data refreshed daily, snapshot 2026-10-07.