Molmo2-CapTest: leaderboard

Metric: Caption F1 (%; LLM-judged statement precision and recall against human captions, 693 videos). Source: arxiv.org. Saturation forecast: Around 2029. 21 models tracked.

Top models

#ModelScore
1GPT-5 Mini56.6
2GPT-550.1
3Gemini 2.5 Flash46
4Molmo2-8B43.2
5Gemini 2.5 Pro42.1
6Gemini 3 Pro36
7Qwen 3 VL 8B26.7
8Claude Sonnet 4.526
9Qwen 3 VL 4B25.2
10GLM-4.1V-9B (Thinking)18.4

Interactive version: theaggregate.ai/benchmark?slug=molmo2-captest · How It Works · Data refreshed daily, snapshot 2026-09-25.