EgoMonth: leaderboard

Metric: Macro-average accuracy (%; mean of the 14 task accuracies over 1,443 four-option multiple-choice questions written by annotators over month-long first-person recordings of 20 participants (over 300 hours, 20 to 120 days each), answers matched to the key without an LLM judge; each model at its official frame setting (16 to 512 frames, 1 fps for Gemini); chance 25%). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro71.8
2Qwen 2 VL 7B54.5

Interactive version: theaggregate.ai/benchmark?slug=egomonth · How It Works · Data refreshed daily, snapshot 2026-09-29.