OmniBehavior: leaderboard

Metric: Overall score, the paper's unweighted mean of its six behaviour metrics (four binary-behaviour F1 scores, a watch-duration score and a judged customer-service dialogue score) for predicting the real behaviour of 200 Kuaishou users (three months of interleaved video, live-stream, ads and e-commerce traces, 8,143 actions per user on average) from the user profile, history and current scenario, 6,000 balanced tasks, 32k context window, temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.544.55
2GLM-4.741.46
3Claude Sonnet 4.540.49
4GPT-5.239.07
5Kimi K2 090537.61
6Claude Sonnet 436.87
7Claude Haiku 4.536.48
8GPT-4o36.27
9Gemini 3 Flash32.6
10Qwen 3 235B A22B 2507 Instruct32.11

Interactive version: theaggregate.ai/benchmark?slug=omnibehavior · How It Works · Data refreshed daily, snapshot 2026-10-07.