PushupBench - R2: leaderboard

Metric: Coefficient of determination R2 between predicted and true repetition counts (at most 1; 0 matches predicting the mean, negative values mean constant or random output; predictions off by more than 50 excluded), on PushupBench's 446 long-form exercise clips (22-117 s, 11 creators in 10 countries; 27 clips accept several valid counts) sampled at 5 fps up to 112 frames (Claude models capped at 100 frames by the API), 360p frames, greedy decoding, 10 prompt templates; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash0.82
2Gemini 3 Pro0.7
3Gemini 2.5 Pro0.63
4Gemini 2.5 Flash0.49
5Qwen 3 VL 32B Instruct0.36
6Qwen 3 VL 32B (Thinking)0.35
7Gemini 2.0 Flash0.25
8Qwen 3 VL 235B A22B Instruct0.17
9Claude Opus 4.50.09
10Qwen 3 VL 30B A3B (Thinking)0.08
11GPT-50.01
12Claude Sonnet 4.50
13Qwen 3 VL 30B A3B Instruct-0.06
14Qwen 3 VL 8B Instruct-0.11
15Llama 4 Maverick-0.15

Interactive version: theaggregate.ai/benchmark?slug=pushupbench-r2 · How It Works · Data refreshed daily, snapshot 2026-10-07.