V-STaR — leaderboard
Spatio-temporal reasoning benchmark for Video-LLMs, reporting category-level modified arithmetic mean (mAM) and modified logarithmic geometric mean (mLGM) over video question chains.
Metric: Average All Score. Source: huggingface.co. Status: years away from saturation. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 32.45 |
| 2 | Gemini 2.0 Flash | 31.22 |
Interactive version: theaggregate.ai/benchmark?slug=v-star · How the rankings work · Data refreshed daily, snapshot 2026-07-22.