Drive-P2D: leaderboard
Metric: Average score (%; unweighted mean of the six task accuracies, zero-shot, all scenarios; 6,650 questions on front-camera images from nuScenes, KITTI and BDD100K, mean of two runs). Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 | 64.84 |
| 2 | Gemini 2.5 Pro | 57.36 |
| 3 | Qwen 2.5 VL 72B Instruct | 57.33 |
| 4 | Llama 3.2 90B Vision Instruct | 56.35 |
| 5 | Phi-4 Multimodal Instruct | 51.89 |
| 6 | Qwen 2.5 VL 7B Instruct | 50 |
| 7 | Llama 3.2 11B Instruct | 49.55 |
| 8 | Gemma 3 27B (IT) | 46.78 |
| 9 | Gemma 3 4B (IT) | 34.1 |
Interactive version: theaggregate.ai/benchmark?slug=drive-p2d · How It Works · Data refreshed daily, snapshot 2026-09-26.