NextMotionQA - Captioning: leaderboard

Metric: Task 2 captioning score (0-100): Qwen3.6-Plus judge rating of content coverage and action consistency of free-form captions of 396 rendered AMASS human-motion clips given to the model as RGB video, mean of the easy, medium and hard subsets; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 12 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B35.57
2Qwen 3.6 Plus35.23
3Qwen 3.5 4B34.87
4Qwen 3.5 9B33.73
5GPT-5.4 Mini32.77
6InternVL3.5-8B32.03
7Qwen 3.5 0.8B23.3

Interactive version: theaggregate.ai/benchmark?slug=nextmotionqa-captioning · How It Works · Data refreshed daily, snapshot 2026-09-29.