V2X-QA - Cooperative Average: leaderboard

Metric: Accuracy (%) on the cooperative view (vehicle and roadside images together), mean of its four tasks; four-option multiple-choice questions on the expert-verified test split, answered as a single option letter under the paper's unified prompt; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.

Top models

#ModelScore
1Qwen 2.5 VL 72B Instruct42
2GPT-5.236.7
3Gemini 3 Flash (Preview)36.7
4Qwen 3.5 Plus34
5Gemini 2.5 Pro30.9
6Qwen 3 VL 30B A3B Instruct27.5
7GPT-5 Mini23.7

Interactive version: theaggregate.ai/benchmark?slug=v2x-qa-cooperative-average · How It Works · Data refreshed daily, snapshot 2026-10-07.