InterLV-Search - Level 3 Multi-Branch: leaderboard
Metric: Level 3 open-web interleaved search, the 340 multi-branch questions (at most 10 interactions) final-answer accuracy judged by GPT-5.4-mini for semantic equivalence, %, through the InterLV-Agent reason-act-observe framework with image search, reverse image search, web search, browsing, cropping and code tools; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 37.94 |
| 2 | GPT-5.4 | 33.82 |
| 3 | Claude Sonnet 4.6 | 33.18 |
| 4 | Qwen 3.6 Plus | 29.41 |
| 5 | GPT-5 | 27.65 |
Interactive version: theaggregate.ai/benchmark?slug=interlv-search-level-3-multi-branch · How It Works · Data refreshed daily, snapshot 2026-10-07.