SaaS-Bench - Text-Only: leaderboard

Metric: Checkpoint score (%) on the 74 text-only tasks of the four text-only domains, weighted-checkpoint score over the final application states, browser-use agent loop over live self-hosted SaaS apps (DOM and screenshot observations, restricted browser actions, no JavaScript or backend access), same step budget, timeout and failure cap for every model; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.742.8
2GPT-5.5 (High)42.1
3Claude Opus 4.640.7
4GPT-5.4 (High)33
5Kimi K2.630.1
6Kimi K2.523.6
7Qwen 3.6 Plus23.1
8DeepSeek V4 Pro21.5
9Gemini 3.1 Pro (Preview)20.6
10Seed 2.0 Pro19.8
11Claude Sonnet 4.618.7
12Gemini 3.5 Flash (High)17.6
13GLM-5.117.4
14MiniMax-M2.715.8

Interactive version: theaggregate.ai/benchmark?slug=saas-bench-text-only · How It Works · Data refreshed daily, snapshot 2026-10-07.