V-STaR — leaderboard

Spatio-temporal reasoning benchmark for Video-LLMs, reporting category-level modified arithmetic mean (mAM) and modified logarithmic geometric mean (mLGM) over video question chains.

Metric: Average All Score. Source: huggingface.co. Status: years away from saturation. 14 models tracked.

Top models

#ModelScore
1GPT-4o32.45
2Gemini 2.0 Flash31.22

Interactive version: theaggregate.ai/benchmark?slug=v-star · How the rankings work · Data refreshed daily, snapshot 2026-07-22.