SEED-Bench — leaderboard
Comprehensive multimodal benchmark with 19K questions across 12 dimensions evaluating image and video understanding.
Metric: Avg Score (%). Source: huggingface.co. Status: saturation imminent. 63 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | [InternVL-Chat-V1.2-Plus](https://github.com/OpenGVLab/InternVL) | 70.4 |
| 2 | [Weitu-VL-1.0](https://weitu.ai/) | 69.2 |
| 3 | [SPHINXv2-1k](https://github.com/Alpha-VLLM/LLaMA2-Accessory/tree/main/SPHINX) | 67.5 |
| 4 | [GPT-4V](https://openai.com/research/gpt-4v-system-card) | 67.3 |
| 5 | [Qwen-VL-plus](https://github.com/QwenLM/Qwen-VL/tree/master?tab=readme-ov-file#qwen-vl-plus) | 66.8 |
| 6 | [SPHINXv1-1k](https://github.com/Alpha-VLLM/LLaMA2-Accessory/tree/main/SPHINX) | 63.9 |
| 7 | [LLaVA-v1.5-LoRA](https://llava-vl.github.io) | 62.8 |
| 8 | [llava-v1.5-7b-finetune]() | 62.8 |
| 9 | [LLaVA-v1.5-13B-LoRA](https://llava-vl.github.io) | 62.4 |
| 10 | [InfMLLM-13B](https://github.com/mightyzau/InfMLLM) | 62.3 |
Interactive version: theaggregate.ai/benchmark?slug=seed-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.