EgoBench - Dynamic Hard Mode: leaderboard
Metric: Joint Success Rate (%; every required tool call issued and the final database state correct, impatient simulated user adding unrelated chatter). Source: arxiv.org. Saturation forecast: Around May 2028. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 16.75 |
| 2 | Qwen 3.5 397B A17B (Non-reasoning) | 14.35 |
| 3 | Qwen 3.6 Plus (Non-reasoning) | 13.88 |
| 4 | Seed 2.0 Pro (Non-reasoning) | 11.58 |
| 5 | Kimi K2.5 (Non-reasoning) | 11.2 |
| 6 | GLM-5V Turbo | 3.33 |
| 7 | MiMo-V2-Omni | 2.3 |
| 8 | Qwen 3 VL 235B A22B Instruct | 0.29 |
Interactive version: theaggregate.ai/benchmark?slug=egobench-dynamic-hard-mode · How It Works · Data refreshed daily, snapshot 2026-09-25.