Omni2Web - Instruction Recovery Action F1: leaderboard

Metric: Action F1 (%; Track B component: recovered edit actions matched one-to-one to the reference actions; 918 bilingual screen recordings of weakly referential webpage edit requests; the model rewrites the recording into an ordered list of explicit target and action instructions without the source HTML; one-to-one matching against reference instructions, semantic equivalence judged by Gemini 3.1 Pro; unusable outputs score zero). Source: arxiv.org. Saturation forecast: Around 2031. 17 models tracked.

Top models

#ModelScore
1Claude Opus 549.76
2GPT-5.6 Sol49.17
3Qwen 3.5 Omni Plus48.29
4Gemini 3.5 Flash47.88
5Grok 4.647.06
6Kimi K345.25
7Seed 2.0 Lite45.14
8Seed 2.1 Pro44.96
9Gemini 3.1 Pro (Preview)44.18
10Qwen 3.8 Max41.9
11MiMo-V2.540.67
12Qwen 3.8 27B39.85
13MiniMax-M337.38
14Muse Spark 1.132.37
15MiniCPM-o-4.531.93

Interactive version: theaggregate.ai/benchmark?slug=omni2web-instruction-recovery-action-f1 · How It Works · Data refreshed daily, snapshot 2026-09-26.