DriveSpatial - Reasoning: leaderboard

Metric: Accuracy (%) on the spatiotemporal reasoning tasks (interactions and future dynamics) on the full benchmark (15,670 human-verified questions over 20 tasks from nuScenes, Waymo, TruckScenes, AV2 and ONCE); open-source generalist and driving or spatial specialist VLMs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScore
1Gemma 3 12B (IT)40.8
2InternVL3-8B30.65

Interactive version: theaggregate.ai/benchmark?slug=drivespatial-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.