MLLM-as-a-Judge — leaderboard

Evaluates multimodal LLMs as judges for scoring and ranking outputs, testing meta-evaluation capabilities across vision-language tasks.

Metric: Average Score. Source: mllm-judge.github.io. 16 models tracked.

Top models

#ModelScore
10.6070.9
20.8040.9
30.6960.8
40.4930.6
50.6570.6
60.4540.5
7LLaVA-1.5-13b0.5
8LLaVA-1.5-13b0.5
9LLaVA-1.5-13b0.5
100.4490.5

Interactive version: theaggregate.ai/benchmark?slug=mllm-as-a-judge · How the rankings work · Data refreshed daily, snapshot 2026-07-22.