VisIT-Bench Multiple Images — leaderboard

VisIT-Bench Multiple Images ranks vision-language models with human-preference Elo scores on instruction-following tasks over multiple images.

Metric: Elo. Source: huggingface.co. 4 models tracked.

Top models

#ModelScore
1Human Verified GPT-4 Reference1192
2mPLUG-Owl995
3Otter911
4OpenFlamingo902

Interactive version: theaggregate.ai/benchmark?slug=visit-bench-multiple-images · How the rankings work · Data refreshed daily, snapshot 2026-07-22.