NextMotionQA - Captioning: leaderboard
Metric: Task 2 captioning score (0-100): Qwen3.6-Plus judge rating of content coverage and action consistency of free-form captions of 396 rendered AMASS human-motion clips given to the model as RGB video, mean of the easy, medium and hard subsets; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 27B | 35.57 |
| 2 | Qwen 3.6 Plus | 35.23 |
| 3 | Qwen 3.5 4B | 34.87 |
| 4 | Qwen 3.5 9B | 33.73 |
| 5 | GPT-5.4 Mini | 32.77 |
| 6 | InternVL3.5-8B | 32.03 |
| 7 | Qwen 3.5 0.8B | 23.3 |
Interactive version: theaggregate.ai/benchmark?slug=nextmotionqa-captioning · How It Works · Data refreshed daily, snapshot 2026-09-29.