Drive-P2D: leaderboard

Metric: Average score (%; unweighted mean of the six task accuracies, zero-shot, all scenarios; 6,650 questions on front-camera images from nuScenes, KITTI and BDD100K, mean of two runs). Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GPT-4.164.84
2Gemini 2.5 Pro57.36
3Qwen 2.5 VL 72B Instruct57.33
4Llama 3.2 90B Vision Instruct56.35
5Phi-4 Multimodal Instruct51.89
6Qwen 2.5 VL 7B Instruct50
7Llama 3.2 11B Instruct49.55
8Gemma 3 27B (IT)46.78
9Gemma 3 4B (IT)34.1

Interactive version: theaggregate.ai/benchmark?slug=drive-p2d · How It Works · Data refreshed daily, snapshot 2026-09-26.