VisIT-Bench Multiple Images — leaderboard
VisIT-Bench Multiple Images ranks vision-language models with human-preference Elo scores on instruction-following tasks over multiple images.
Metric: Elo. Source: huggingface.co. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Human Verified GPT-4 Reference | 1192 |
| 2 | mPLUG-Owl | 995 |
| 3 | Otter | 911 |
| 4 | OpenFlamingo | 902 |
Interactive version: theaggregate.ai/benchmark?slug=visit-bench-multiple-images · How the rankings work · Data refreshed daily, snapshot 2026-07-22.