Omni2Web - Instruction Recovery: leaderboard

Metric: Instruction Recovery Score (%; Track B: mean of target F1 and action F1 per instance, macro-averaged over 918 bilingual screen recordings of weakly referential webpage edit requests; the model rewrites the recording into an ordered list of explicit target and action instructions without the source HTML; one-to-one matching against reference instructions, semantic equivalence judged by Gemini 3.1 Pro; unusable outputs score zero). Source: arxiv.org. Saturation forecast: Around 2030. 17 models tracked.

Top models

#ModelScore
1Qwen 3.5 Omni Plus49.14
2Claude Opus 548.91
3GPT-5.6 Sol48.34
4Gemini 3.5 Flash47.46
5Seed 2.0 Lite46.52
6Gemini 3.1 Pro (Preview)43.19
7Kimi K341.95
8Grok 4.640.83
9Seed 2.1 Pro39.08
10Qwen 3.8 Max38.37
11MiMo-V2.537.31
12Qwen 3.8 27B29.84
13MiniMax-M329.32
14Muse Spark 1.125.41
15MiniCPM-o-4.524.9

Interactive version: theaggregate.ai/benchmark?slug=omni2web-instruction-recovery · How It Works · Data refreshed daily, snapshot 2026-09-26.