DriveSpatial - Reasoning: leaderboard
Metric: Accuracy (%) on the spatiotemporal reasoning tasks (interactions and future dynamics) on the full benchmark (15,670 human-verified questions over 20 tasks from nuScenes, Waymo, TruckScenes, AV2 and ONCE); open-source generalist and driving or spatial specialist VLMs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 3 12B (IT) | 40.8 |
| 2 | InternVL3-8B | 30.65 |
Interactive version: theaggregate.ai/benchmark?slug=drivespatial-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.