PaveVQA - Localization Token-F1: leaderboard

Metric: Token-level F1 (%) of short spatial descriptions on localization questions, of the zero-shot model on PaveVQA, the vision-language question answering part of PaveBench (real highway pavement images, single-turn, multi-turn and expert-corrected questions); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1DeepSeek-VL2-small40.36
2LLaVA-OneVision-7B27.71
3Qwen2.5-VL-3B16.39

Interactive version: theaggregate.ai/benchmark?slug=pavevqa-localization-token-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.