E3VS-Bench - Context-Guided Search: leaderboard
Metric: Judge score on Context-guided Search questions (finding an object from its function or situation without its name) of the E3VS-Bench test split (377 episodes, 21 3D Gaussian Splatting scenes): the model drives a 5-DoF camera agent (at most 25 steps) to a viewpoint from which the question can be answered; GPT-5.1 then answers from the agent's final view and a GPT-5.1 judge scores the answer 5 (correct) or 1 (incorrect); mean judge score on a 1-5 scale; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 3.4 |
| 2 | Step3 VL 10B | 3.13 |
| 3 | GPT-5.1 (Non-reasoning) | 2.73 |
| 4 | Gemini 2.5 Pro | 2.6 |
Interactive version: theaggregate.ai/benchmark?slug=e3vs-bench-context-guided-search · How It Works · Data refreshed daily, snapshot 2026-10-07.