VIABench - VQA: leaderboard
Metric: Visual question answering score (0-100): answers to questions about the environment or objects asked at a timestamp of a blind user video, given the video up to that point, judged by GPT-5 as the semantic similarity (0 to 1, reported times 100) between the model answer and the ground-truth annotation, on first-person videos recorded or shared by visually impaired people, with each model at its default frame sampling; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 62.9 |
| 2 | Gemini 2.5 Pro | 58.8 |
| 3 | GPT-4o (2024-08-06) | 56.6 |
| 4 | InternVL3.5-8B | 39.5 |
| 5 | Qwen 2.5 VL 7B | 38.9 |
Interactive version: theaggregate.ai/benchmark?slug=viabench-vqa · How It Works · Data refreshed daily, snapshot 2026-09-29.