Drive-P2D - Special Scene Factors (Scene-2): leaderboard
Metric: Accuracy (%; identify special scene factors such as roadworks or accidents that constrain the decision (multiple choice, partial credit for a correct subset); zero-shot, all scenarios, mean of two runs). Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 | 46.44 |
| 2 | Qwen 2.5 VL 72B Instruct | 37.62 |
| 3 | Llama 3.2 90B Vision Instruct | 36.78 |
| 4 | Qwen 2.5 VL 7B Instruct | 33.6 |
| 5 | Gemini 2.5 Pro | 25.61 |
| 6 | Phi-4 Multimodal Instruct | 21.67 |
| 7 | Llama 3.2 11B Instruct | 18.37 |
| 8 | Gemma 3 27B (IT) | 7.2 |
| 9 | Gemma 3 4B (IT) | 0.89 |
Interactive version: theaggregate.ai/benchmark?slug=drive-p2d-special-scene-factors-scene-2 · How It Works · Data refreshed daily, snapshot 2026-09-26.