EgoIntrospect - Valuable Interaction Identification: leaderboard

Metric: Accuracy (%) on 1,161 four-choice questions asking which of four proactive assistant requests the wearer accepted for a clip (distractors are requests the wearer did not accept), fixation-overlaid egocentric clips from 15 held-out participants with the participant profile; native audio-video models hear the clip audio, the other models read its transcript; temperature 0 (0.6 for InternVL3.5 thinking mode); higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)55.81
2Kimi K2.554.78
3Gemini 3 Flash (Preview)52.8
4Kimi K2.5 (Thinking)52.37
5Qwen 3.5 397B A17B51.42
6Qwen 3.5 Omni Plus49.96
7Qwen 3 VL 8B Instruct47.63
8Qwen3 Omni 30B A3B Instruct46.34
9Qwen 3 VL 8B (Thinking)46.17
10GPT-4o43.84
11InternVL3-8B42.89
12Qwen 2.5 VL 7B Instruct41.09

Interactive version: theaggregate.ai/benchmark?slug=egointrospect-valuable-interaction-identification · How It Works · Data refreshed daily, snapshot 2026-10-07.