Driving View Evidence QA: leaderboard
Metric: View-selection exact-match accuracy (%), mean of 5 runs: pick the camera that holds the evidence (or None when no view supports the question) on 122 conflict-centric questions over six synchronized nuScenes camera views from 73 scenes (causality, counterfactual and intent), zero-shot, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 77.54 |
| 2 | Qwen 2.5 VL 7B Instruct | 12.62 |
Interactive version: theaggregate.ai/benchmark?slug=driving-view-evidence-qa · How It Works · Data refreshed daily, snapshot 2026-09-29.