MA-Bench - Coarse-Grained Recognition: leaderboard

Metric: Accuracy (%) on the coarse-grained (body-level) micro-action recognition questions (random 14.7), MA-Bench's 1,000 micro-action videos (about 2 seconds each, 52 micro-action categories), 8 sampled frames, zero-shot, final answer without intermediate reasoning; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 23 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash43#237
2Qwen 2 VL 7B28.5#741
3GPT-4o20.5#333
4Qwen 2.5 VL 7B Instruct20.4#643
5InternVL2-8B17.4#826
6InternVL3-8B17.1#606
7Pixtral-12B14.4#795
8Phi-4 Multimodal Instruct11.4#896

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=ma-bench-coarse-grained-recognition · How It Works · Data refreshed daily, snapshot 2026-10-11.