MacAgentBench (Agent-S3) - Multi-App: leaderboard

Metric: Pass@1 (%; share of the 140 Multi-App tasks solved on a single attempt, deterministic rule-based checks; the Agent S3 multi-agent framework, at most 50 steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Agent S3 + GPT-5.441.4
2Agent S3 + Claude Opus 4.640.7
3Agent S3 + Gemini 3.1 Pro30.7
4Agent S3 + Qwen3-VL-235B-A22B25.7
5Agent S3 + Qwen3-VL-32B-Thinking19.3
6Agent S3 + Qwen3-VL-8B-Thinking12.1
7Agent S3 + InternVL3.5-14B1.4
8Agent S3 + InternVL3.5-8B0.7

Interactive version: theaggregate.ai/benchmark?slug=macagentbench-agent-s3-multi-app · How It Works · Data refreshed daily, snapshot 2026-09-26.