Omni2Web - Direct Editing: leaderboard

Metric: Edit Fidelity Score (%; Track A: 0.10 target grounding + 0.50 edit fulfillment + 0.20 no-damage + 0.20 revision handling per instance, averaged over 918 bilingual screen recordings of weakly referential webpage edit requests (speech plus cursor), 13,907 edit steps; the model edits the source HTML directly from the recording under its native input interface (omni: audio-visual video; vision-language: video or 1 fps frames plus manual transcript); rubric items judged by Gemini 3.1 Pro; unusable outputs score zero). Source: arxiv.org. Saturation forecast: Around 2028. 15 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash51.17
2Claude Opus 547.37
3GPT-5.6 Sol42.86
4Grok 4.641.45
5Kimi K340.63
6Qwen 3.8 Max38.6
7Seed 2.1 Pro34.48
8Gemini 3.1 Pro (Preview)34.08
9Qwen 3.5 Omni Plus29.56
10MiniMax-M329.49
11MiMo-V2.528.79
12Seed 2.0 Lite28.33
13Qwen 3.8 27B26.94
14Muse Spark 1.124.08

Interactive version: theaggregate.ai/benchmark?slug=omni2web-direct-editing · How It Works · Data refreshed daily, snapshot 2026-09-26.