VisIT-Bench Single Image — leaderboard
VisIT-Bench Single Image ranks vision-language models with human-preference Elo scores on instruction-following tasks over single images.
Metric: Elo. Source: huggingface.co. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Human Verified GPT-4 Reference | 1370 |
| 2 | LLaVA (13B) | 1106 |
| 3 | LlamaAdapter-v2 (7B) | 1082 |
| 4 | mPLUG-Owl (7B) | 1081 |
| 5 | InstructBLIP (13B) | 1011 |
| 6 | Otter (9B) | 991 |
| 7 | VisualGPT (Da Vinci 003) | 972 |
| 8 | MiniGPT-4 (7B) | 921 |
| 9 | OpenFlamingo (9B) | 877 |
| 10 | PandaGPT (13B) | 826 |
Interactive version: theaggregate.ai/benchmark?slug=visit-bench-single-image · How the rankings work · Data refreshed daily, snapshot 2026-07-22.