InfiMM-Eval — leaderboard
Metric: Overall score. Source: huggingface.co. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4V | 74.44 |
| 2 | SPHINX v2 | 39.48 |
| 3 | Qwen-VL-Chat | 37.39 |
| 4 | CogVLM-Chat | 37.16 |
| 5 | LLaVA-1.5 | 32.62 |
| 6 | LLaMA-Adapter V2 | 30.46 |
| 7 | Emu | 28.24 |
| 8 | InstructBLIP | 28.02 |
| 9 | InternLM-XComposer-VL | 26.84 |
| 10 | Otter | 22.69 |
Interactive version: theaggregate.ai/benchmark?slug=infimm-eval · How the rankings work · Data refreshed daily, snapshot 2026-07-22.