MERRIN (No Search): leaderboard

Metric: Accuracy (%) with no search or browsing tool, on MERRIN's 162 human-annotated questions with unique short answers that need non-text web evidence (image, video, audio or table) and multi-hop reasoning or conflict resolution across noisy sources; accuracy against the gold answer by an LLM judge with the BrowseComp prompt, mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)24.7
2Gemini 3 Pro23.5
3Gemini 3 Flash19.1
4GPT-5.4 Mini14
5Gemini 3.1 Flash Lite12.8
6Qwen 3 235B A22B12.1
7Qwen 3 4B10.3
8GPT-5.4 Nano9.9
9Qwen 3 30B A3B8

Interactive version: theaggregate.ai/benchmark?slug=merrin-no-search · How It Works · Data refreshed daily, snapshot 2026-10-07.