EgoMonth: leaderboard
Metric: Macro-average accuracy (%; mean of the 14 task accuracies over 1,443 four-option multiple-choice questions written by annotators over month-long first-person recordings of 20 participants (over 300 hours, 20 to 120 days each), answers matched to the key without an LLM judge; each model at its official frame setting (16 to 512 frames, 1 fps for Gemini); chance 25%). Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 71.8 |
| 2 | Qwen 2 VL 7B | 54.5 |
Interactive version: theaggregate.ai/benchmark?slug=egomonth · How It Works · Data refreshed daily, snapshot 2026-09-29.