Omni2Web - Instruction Utility: leaderboard
Metric: Edit Fidelity Score (%; Track C: the model's Track B instructions are executed on the source HTML by a fixed DeepSeek-v4-pro coding model, then scored with the Track A rubric: 0.10 grounding + 0.50 edit fulfillment + 0.20 no-damage + 0.20 revision handling; 918 instances; judged by Gemini 3.1 Pro). Source: arxiv.org. Saturation forecast: Around 2031. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 Omni Plus | 51.08 |
| 2 | Gemini 3.5 Flash | 49.77 |
| 3 | Seed 2.0 Lite | 47.42 |
| 4 | Claude Opus 5 | 47.07 |
| 5 | GPT-5.6 Sol | 46.19 |
| 6 | Gemini 3.1 Pro (Preview) | 45.53 |
| 7 | Kimi K3 | 40.91 |
| 8 | Qwen 3.8 Max | 39.52 |
| 9 | Seed 2.1 Pro | 38.66 |
| 10 | Grok 4.6 | 38.6 |
| 11 | MiMo-V2.5 | 37.87 |
| 12 | MiniCPM-o-4.5 | 32.45 |
| 13 | MiniMax-M3 | 28.67 |
| 14 | Qwen 3.8 27B | 28.48 |
| 15 | Muse Spark 1.1 | 24.37 |
Interactive version: theaggregate.ai/benchmark?slug=omni2web-instruction-utility · How It Works · Data refreshed daily, snapshot 2026-09-26.