EgoBench - Dynamic Hard Mode - Micro Accuracy: leaderboard

Metric: MicroAcc (%; share of ground-truth tool calls issued, pooled over tasks, impatient simulated user adding unrelated chatter). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)42.89
2Seed 2.0 Pro (Non-reasoning)37.22
3Qwen 3.5 397B A17B (Non-reasoning)36.97
4Qwen 3.6 Plus (Non-reasoning)32.95
5Kimi K2.5 (Non-reasoning)31.01
6GLM-5V Turbo22.41
7MiMo-V2-Omni15.35
8Qwen 3 VL 235B A22B Instruct4.17

Interactive version: theaggregate.ai/benchmark?slug=egobench-dynamic-hard-mode-micro-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-25.