MVX-Bench - Re-Identification: leaderboard

Metric: Accuracy (%) on the 150 Re-Identification questions (decide whether two camera views show the same person), MVX-Bench's multi-video multiple-choice questions (each pairs a question with several videos; 1,442 questions over 4,255 videos repurposed from 11 annotated video datasets), end-to-end inference at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 21 models tracked.

Top models

#ModelScoreOverall rank
1Gemma 3 12B (IT)60#655
2Qwen 3.6 27B (Thinking)60#163 (Qwen 3.6 27B)
3Qwen 2.5 VL 7B Instruct58.7#643
4GPT-5.258#105
5Pixtral-12B58#795
6Qwen 3.6 27B (Non-reasoning)57.3#163 (Qwen 3.6 27B)
7Gemma 3 27B (IT)56.7#509
8Qwen 2.5 VL 32B Instruct56.7#443
9Phi-4 Multimodal Instruct55.3#896
10Gemma 3 4B (IT)54#971
11GPT-4o40#333

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mvx-bench-re-identification · How It Works · Data refreshed daily, snapshot 2026-10-11.