VistaHop (No Tools): leaderboard

Metric: Pass@1 (%) on 600 Visual DeepSearch tasks (25 scenarios; 154 with 5 to 9 evidence steps and 446 with 10 or more) that require repeatedly inspecting image regions and chaining visual and web evidence, run in the VistaArena agent loop (at most 10 rounds, one tool call per round, temperature 0.7, five runs averaged) with a Qwen3-VL-32B-Instruct judge, without tools (a single answer from the image and query); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1GPT-5.211.33
2Claude Sonnet 4.511.08
3Gemini 2.5 Pro8.63
4Qwen 3 VL 235B A22B Instruct7.84
5Qwen 3 VL 30B A3B Instruct6.72

Interactive version: theaggregate.ai/benchmark?slug=vistahop-no-tools · How It Works · Data refreshed daily, snapshot 2026-09-29.