VisIT-Bench Single Image — leaderboard

VisIT-Bench Single Image ranks vision-language models with human-preference Elo scores on instruction-following tasks over single images.

Metric: Elo. Source: huggingface.co. 11 models tracked.

Top models

#ModelScore
1Human Verified GPT-4 Reference1370
2LLaVA (13B)1106
3LlamaAdapter-v2 (7B)1082
4mPLUG-Owl (7B)1081
5InstructBLIP (13B)1011
6Otter (9B)991
7VisualGPT (Da Vinci 003)972
8MiniGPT-4 (7B)921
9OpenFlamingo (9B)877
10PandaGPT (13B)826

Interactive version: theaggregate.ai/benchmark?slug=visit-bench-single-image · How the rankings work · Data refreshed daily, snapshot 2026-07-22.