SaaSBench (Engineering): leaderboard

Metric: Task score (%; pass@1, one rollout per task: mean over the 30 tasks of the share of each task's validation-DAG max score earned, with deterministic HTTP, database, container and browser checks plus Claude Sonnet 4.5 rubric nodes; the agent builds and deploys an enterprise SaaS system from a PRD in a Docker workspace, 10,800 s budget). Source: arxiv.org. Saturation forecast: Around July 2028. 16 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (Claude Code)20.68

Interactive version: theaggregate.ai/benchmark?slug=saasbench-engineering · How It Works · Data refreshed daily, snapshot 2026-09-26.