SaaS-Bench - Multimodal: leaderboard

Metric: Checkpoint score (%) on the 32 multimodal tasks (images, documents and media files as inputs), weighted-checkpoint score over the final application states, browser-use agent loop over live self-hosted SaaS apps (DOM and screenshot observations, restricted browser actions, no JavaScript or backend access), same step budget, timeout and failure cap for every model; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.648.7
2GPT-5.5 (High)47.6
3Claude Opus 4.746.3
4GPT-5.4 (High)46.1
5Qwen 3.6 Plus45.5
6Seed 2.0 Pro44.2
7Kimi K2.643.5
8Gemini 3.1 Pro (Preview)42.4
9Kimi K2.537
10Gemini 3.5 Flash (High)36.5
11Claude Sonnet 4.633.9

Interactive version: theaggregate.ai/benchmark?slug=saas-bench-multimodal · How It Works · Data refreshed daily, snapshot 2026-10-07.