VisIT-Bench Multiple Images: leaderboard
VisIT-Bench Multiple Images ranks vision-language models with human-preference Elo scores on instruction-following tasks over multiple images.
Metric: Elo. Source: huggingface.co. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Human Verified GPT-4 Reference | 1192 |
| 2 | mPLUG-Owl | 995 |
| 3 | Otter | 911 |
| 4 | OpenFlamingo | 902 |
Interactive version: theaggregate.ai/benchmark?slug=visit-bench-multiple-images · How It Works · Data refreshed daily, snapshot 2026-09-05.