EgoIntrospect - Proactive Timing Judgment: leaderboard

Metric: Macro-F1 (%) over the interruptible and do-not-disturb classes of a binary judgment of whether a proactive assistant may interrupt the wearer during a clip (269 clips, labels from the wearers' own do-not-disturb annotations, interruptible the majority class), fixation-overlaid egocentric clips from 15 held-out participants with the participant profile; native audio-video models hear the clip audio, the other models read its transcript; temperature 0 (0.6 for InternVL3.5 thinking mode); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 21 models tracked.

Top models

#ModelScore
1Qwen 3 VL 8B Instruct60.69
2Kimi K2.559.78
3Qwen 3.5 397B A17B58.36
4Qwen3 Omni 30B A3B Instruct57.62
5Qwen 3 VL 8B (Thinking)57.57
6Gemini 3 Flash (Preview)57.41
7Gemini 3.1 Pro (Preview)57.13
8InternVL3-8B56.22
9Kimi K2.5 (Thinking)54.21
10Qwen 3.5 Omni Plus53.22
11GPT-4o29.55
12Qwen 2.5 VL 7B Instruct28.27

Interactive version: theaggregate.ai/benchmark?slug=egointrospect-proactive-timing-judgment · How It Works · Data refreshed daily, snapshot 2026-10-07.