AffectSim (Active Observation): leaderboard

Metric: Macro-F1 (%; A-Obs: the view reached by the authors' two-stage active-observation baseline (person search, then target following), falling back to P-Init when no new view is acquired; five-class emotion recognition (angry, fearful, happy, neutral, sad) on the 3,941 test episodes of AffectSim, full-body affective motions replayed in Habitat indoor scenes; frozen recognizer, eight uniformly sampled frames for most open models, full video or eight frames for closed models). Source: arxiv.org. Saturation forecast: Around 2030. 24 models tracked.

Top models

#ModelScore
1Claude Opus 528.37
2Gemini 3 Pro25.42
3Gemini 3 Flash24.62
4GPT-5.6 Sol21.87
5GPT-5.221
6InternVL2.5-78B18.52
7Qwen 3 VL 8B Instruct13.19
8Qwen 3 VL 8B (Thinking)9.59
9Qwen2.5-Omni-7B9.45
10Qwen 2.5 VL 72B Instruct7.63
11MiniCPM-V-2.66.79
12Qwen 2.5 VL 7B Instruct6.11

Interactive version: theaggregate.ai/benchmark?slug=affectsim-active-observation · How It Works · Data refreshed daily, snapshot 2026-09-29.