MMHBench - Third-Person: leaderboard

Metric: Accuracy (%; the 605 third-person assessment multiple-choice questions on 268 long-form mental-health-related videos, 64 uniformly sampled frames, zero-shot about observable behavior and multimodal evidence; higher is better). Source: arxiv.org. Saturation forecast: Around February 2028. 21 models tracked.

Top models

#ModelScore
1GPT-5.563.64
2Qwen 3.7 Plus58.35
3Qwen 3.6 Plus58.18
4Qwen 3.6 35B A3B57.85
5Qwen 3.5 35B A3B55.37
6Qwen 3.5 9B53.22
7Qwen 3.5 4B51.9
8Qwen 3 VL 32B50.74
9Gemma 4 26B A4B45.95
10MiniCPM-V-2.643.31
11Ministral 3 14B30.08
12Gemma 4 E4B28.6

Interactive version: theaggregate.ai/benchmark?slug=mmhbench-third-person · How It Works · Data refreshed daily, snapshot 2026-09-29.