PaveVQA - Localization Token-F1: leaderboard
Metric: Token-level F1 (%) of short spatial descriptions on localization questions, of the zero-shot model on PaveVQA, the vision-language question answering part of PaveBench (real highway pavement images, single-turn, multi-turn and expert-corrected questions); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek-VL2-small | 40.36 |
| 2 | LLaVA-OneVision-7B | 27.71 |
| 3 | Qwen2.5-VL-3B | 16.39 |
Interactive version: theaggregate.ai/benchmark?slug=pavevqa-localization-token-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.