MA-Bench (Open-Ended): leaderboard

Metric: GPT-4o judge score (0 to 5) averaged over the six open-ended dimensions: descriptive understanding (core action semantics, spatial accuracy, temporal consistency) and reasoning and explanation (body-level label, action-level label, causal reasoning chain), MA-Bench's 1,000 micro-action videos (about 2 seconds each, 52 micro-action categories), 8 sampled frames, zero-shot, final answer without intermediate reasoning; higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 23 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash0.76#237
2GPT-4o0.73#333
3InternVL2-8B0.52#826
4Qwen 2.5 VL 7B Instruct0.37#643
5Qwen 2 VL 7B0.31#741
6Pixtral-12B0.18#795
7InternVL3-8B0.1#606
8Phi-4 Multimodal Instruct0.1#896

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=ma-bench-open-ended · How It Works · Data refreshed daily, snapshot 2026-10-11.