Driving View Evidence QA: leaderboard

Metric: View-selection exact-match accuracy (%), mean of 5 runs: pick the camera that holds the evidence (or None when no view supports the question) on 122 conflict-centric questions over six synchronized nuScenes camera views from 73 scenes (causality, counterfactual and intent), zero-shot, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1GPT-5.477.54
2Qwen 2.5 VL 7B Instruct12.62

Interactive version: theaggregate.ai/benchmark?slug=driving-view-evidence-qa · How It Works · Data refreshed daily, snapshot 2026-09-29.