OmniBehavior: leaderboard
Metric: Overall score, the paper's unweighted mean of its six behaviour metrics (four binary-behaviour F1 scores, a watch-duration score and a judged customer-service dialogue score) for predicting the real behaviour of 200 Kuaishou users (three months of interleaved video, live-stream, ads and e-commerce traces, 8,143 actions per user on average) from the user profile, history and current scenario, 6,000 balanced tasks, 32k context window, temperature 0.1; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.5 | 44.55 |
| 2 | GLM-4.7 | 41.46 |
| 3 | Claude Sonnet 4.5 | 40.49 |
| 4 | GPT-5.2 | 39.07 |
| 5 | Kimi K2 0905 | 37.61 |
| 6 | Claude Sonnet 4 | 36.87 |
| 7 | Claude Haiku 4.5 | 36.48 |
| 8 | GPT-4o | 36.27 |
| 9 | Gemini 3 Flash | 32.6 |
| 10 | Qwen 3 235B A22B 2507 Instruct | 32.11 |
Interactive version: theaggregate.ai/benchmark?slug=omnibehavior · How It Works · Data refreshed daily, snapshot 2026-10-07.