VistaHop (No Tools): leaderboard
Metric: Pass@1 (%) on 600 Visual DeepSearch tasks (25 scenarios; 154 with 5 to 9 evidence steps and 446 with 10 or more) that require repeatedly inspecting image regions and chaining visual and web evidence, run in the VistaArena agent loop (at most 10 rounds, one tool call per round, temperature 0.7, five runs averaged) with a Qwen3-VL-32B-Instruct judge, without tools (a single answer from the image and query); higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 11.33 |
| 2 | Claude Sonnet 4.5 | 11.08 |
| 3 | Gemini 2.5 Pro | 8.63 |
| 4 | Qwen 3 VL 235B A22B Instruct | 7.84 |
| 5 | Qwen 3 VL 30B A3B Instruct | 6.72 |
Interactive version: theaggregate.ai/benchmark?slug=vistahop-no-tools · How It Works · Data refreshed daily, snapshot 2026-09-29.