EgoCoT-Bench - Hand-Object Temporal Retrospection: leaderboard

Metric: Answer accuracy (%, four-way multiple choice, strict exact match) of zero-shot MLLMs on egocentric manipulation videos of EgoCoT-Bench (351 clips, 3,172 human-verified questions built from spatio-temporal scene graphs), Hand-Object Temporal Retrospection subtask (order earlier hand-object interactions in time); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 19 models tracked.

Top models

#ModelScore
1GPT-5.259.38
2Qwen 3.5 27B59.38
3Qwen 3.5 397B A17B56.25
4GPT-5.156.25
5Qwen 3.5 122B A10B56.25
6Qwen 3 VL 235B A22B56.25
7Qwen 3.5 Plus53.12

Interactive version: theaggregate.ai/benchmark?slug=egocot-bench-hand-object-temporal-retrospection · How It Works · Data refreshed daily, snapshot 2026-10-07.