PushupBench - Mean Absolute Error: leaderboard

Metric: Mean absolute error of the predicted repetition count (repetitions; predictions off by more than 50 are excluded as parsing or hallucination outliers), on PushupBench's 446 long-form exercise clips (22-117 s, 11 creators in 10 countries; 27 clips accept several valid counts) sampled at 5 fps up to 112 frames (Claude models capped at 100 frames by the API), 360p frames, greedy decoding, 10 prompt templates; lower is better. Source: arxiv.org. Saturation forecast: Around May 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash2.9
2Gemini 3 Pro3.6
3Gemini 2.5 Pro5.7
4Qwen 3 VL 235B A22B Instruct7.1
5GPT-57.6
6Gemini 2.5 Flash7.7
7Qwen 3 VL 32B Instruct7.7
8Gemini 2.0 Flash7.8
9Qwen 2.5 VL 7B7.8
10Qwen 3 VL 8B Instruct7.9
11Qwen 3 VL 4B Instruct8.2
12Qwen 3 VL 30B A3B (Thinking)8.8
13Qwen 3 VL 4B (Thinking)8.9
14Claude Sonnet 4.59
15Qwen 3 VL 32B (Thinking)9.1

Interactive version: theaggregate.ai/benchmark?slug=pushupbench-mean-absolute-error · How It Works · Data refreshed daily, snapshot 2026-10-07.