BenchLMM — leaderboard
Metric: GPT-3.5 score. Source: huggingface.co. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4V | 58.37 |
| 2 | Sphinx-V2-1K | 57.43 |
| 3 | LLaVA-1.5-13B | 55.53 |
| 4 | LLaVA-1.5-7B | 46.83 |
| 5 | InstructBLIP-13B | 45.03 |
| 6 | InstructBLIP-7B | 44.63 |
| 7 | LLaVA-1-13B | 43.5 |
| 8 | Otter-7B | 39.13 |
| 9 | MiniGPT4-13B | 34.93 |
| 10 | MiniGPTv2-7B | 30.1 |
Interactive version: theaggregate.ai/benchmark?slug=benchlmm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.