AgentViSS - Interaction Regulation (Verbalized Vision): leaderboard

Metric: Interaction-regulation task score, the agent first verbalizing the visual cues it sees; normalized task score (0-100), each role-task instance labeled Achieved 2, Partially Achieved 0.5 or Not Achieved 0 by majority vote of three judges (Gemini 3.1 Pro Preview, GPT-5.4, Qwen3.5-27B), over 2,340 role-task instances in 240 multi-party social scenarios with group images and role portraits; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.671.75
2Qwen 3.5 122B A10B61.71
3Qwen 3.5 27B56.28
4GLM-4.6V47.39
5Qwen 3.5 9B47.35
6GPT-5.431.37

Interactive version: theaggregate.ai/benchmark?slug=agentviss-interaction-regulation-verbalized-vision · How It Works · Data refreshed daily, snapshot 2026-09-29.