Q-Bench A2 Single — leaderboard
Metric: Sum Score. Source: huggingface.co. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | BlueImage-GPT (Close-Source) | 5.49 |
| 2 | InternLM-XComposer-VL (InternLM) | 4.21 |
| 3 | Emu2-Chat (LLaMA-33B) | 4.19 |
| 4 | Kosmos-2 | 4.03 |
| 5 | mPLUG-Owl (LLaMA-7B) | 3.94 |
| 6 | LLaVA-v1 (Vicuna-13B) | 3.76 |
| 7 | mPLUG-Owl2 (LLaMA-7B) | 3.67 |
| 8 | MiniGPT-4 (Vicuna-13B) | 3.67 |
| 9 | SPHINX | 3.65 |
| 10 | Otter-v1 (MPT-7B) | 3.61 |
Interactive version: theaggregate.ai/benchmark?slug=q-bench-a2-single · How the rankings work · Data refreshed daily, snapshot 2026-07-22.