VIABench - VQA: leaderboard

Metric: Visual question answering score (0-100): answers to questions about the environment or objects asked at a timestamp of a blind user video, given the video up to that point, judged by GPT-5 as the semantic similarity (0 to 1, reported times 100) between the model answer and the ground-truth annotation, on first-person videos recorded or shared by visually impaired people, with each model at its default frame sampling; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1GPT-562.9
2Gemini 2.5 Pro58.8
3GPT-4o (2024-08-06)56.6
4InternVL3.5-8B39.5
5Qwen 2.5 VL 7B38.9

Interactive version: theaggregate.ai/benchmark?slug=viabench-vqa · How It Works · Data refreshed daily, snapshot 2026-09-29.