Q-Bench A2 Single — leaderboard

Metric: Sum Score. Source: huggingface.co. 22 models tracked.

Top models

#ModelScore
1BlueImage-GPT (Close-Source)5.49
2InternLM-XComposer-VL (InternLM)4.21
3Emu2-Chat (LLaMA-33B)4.19
4Kosmos-24.03
5mPLUG-Owl (LLaMA-7B)3.94
6LLaVA-v1 (Vicuna-13B)3.76
7mPLUG-Owl2 (LLaMA-7B)3.67
8MiniGPT-4 (Vicuna-13B)3.67
9SPHINX3.65
10Otter-v1 (MPT-7B)3.61

Interactive version: theaggregate.ai/benchmark?slug=q-bench-a2-single · How the rankings work · Data refreshed daily, snapshot 2026-07-22.