LivingScreen - Closed-Loop Application: leaderboard

Metric: Success rate (%) on L3 fact-checking, content moderation and preference-simulation tasks graded on the final platform state; LivingScreen: 499 Chinese short-video tasks in a browser replica of a short-video app (1,528 real videos, about 5.6 per feed); a single-agent loop issues GUI actions plus watch and wait primitives from screenshots or recorded clips, 30-step cap, temperature 0.6, high reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 13 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash (High)57.4
2Gemini 3.1 Pro (Preview) (High)46.8
3Seed 2.0 Pro (High)37.6
4Claude Opus 4.6 (High)27.7
5GLM-5V Turbo12.9
6Claude Sonnet 4.6 (High)11.3
7Kimi K2.5 (Thinking)10.9
8GPT-5.4 (High)0
9GPT-5.5 (High)0

Interactive version: theaggregate.ai/benchmark?slug=livingscreen-closed-loop-application · How It Works · Data refreshed daily, snapshot 2026-09-29.