AgentViSS - Interaction Outcome (Verbalized Vision): leaderboard
Metric: Interaction-outcome task score (achieving the role's visually grounded social goal), the agent first verbalizing the visual cues it sees; normalized task score (0-100), each role-task instance labeled Achieved 2, Partially Achieved 0.5 or Not Achieved 0 by majority vote of three judges (Gemini 3.1 Pro Preview, GPT-5.4, Qwen3.5-27B), over 2,340 role-task instances in 240 multi-party social scenarios with group images and role portraits; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 90.64 |
| 2 | GPT-5.4 | 79.7 |
| 3 | Qwen 3.5 122B A10B | 77.99 |
| 4 | Qwen 3.5 27B | 77.35 |
| 5 | Qwen 3.5 9B | 72.86 |
| 6 | GLM-4.6V | 70.47 |
Interactive version: theaggregate.ai/benchmark?slug=agentviss-interaction-outcome-verbalized-vision · How It Works · Data refreshed daily, snapshot 2026-09-29.