InterLV-Search - Level 3 Single-Chain: leaderboard

Metric: Level 3 open-web interleaved search, the 521 single-chain questions (at most 10 interactions) final-answer accuracy judged by GPT-5.4-mini for semantic equivalence, %, through the InterLV-Agent reason-act-observe framework with image search, reverse image search, web search, browsing, cropping and code tools; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 8 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)52.02
2GPT-5.451.06
3Claude Sonnet 4.646.83
4Qwen 3.6 Plus42.8
5GPT-540.12

Interactive version: theaggregate.ai/benchmark?slug=interlv-search-level-3-single-chain · How It Works · Data refreshed daily, snapshot 2026-10-07.