Omni2Web - Instruction Utility: leaderboard

Metric: Edit Fidelity Score (%; Track C: the model's Track B instructions are executed on the source HTML by a fixed DeepSeek-v4-pro coding model, then scored with the Track A rubric: 0.10 grounding + 0.50 edit fulfillment + 0.20 no-damage + 0.20 revision handling; 918 instances; judged by Gemini 3.1 Pro). Source: arxiv.org. Saturation forecast: Around 2031. 17 models tracked.

Top models

#ModelScore
1Qwen 3.5 Omni Plus51.08
2Gemini 3.5 Flash49.77
3Seed 2.0 Lite47.42
4Claude Opus 547.07
5GPT-5.6 Sol46.19
6Gemini 3.1 Pro (Preview)45.53
7Kimi K340.91
8Qwen 3.8 Max39.52
9Seed 2.1 Pro38.66
10Grok 4.638.6
11MiMo-V2.537.87
12MiniCPM-o-4.532.45
13MiniMax-M328.67
14Qwen 3.8 27B28.48
15Muse Spark 1.124.37

Interactive version: theaggregate.ai/benchmark?slug=omni2web-instruction-utility · How It Works · Data refreshed daily, snapshot 2026-09-26.