MacAgentBench (Agent-S3): leaderboard

Metric: Pass@1 (%; share of the 676 macOS desktop tasks in 25 applications solved on a single attempt, judged by deterministic rule-based state checks; the Agent S3 multi-agent framework, at most 50 steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Agent S3 + Claude Opus 4.666.9
2Agent S3 + GPT-5.458.9
3Agent S3 + Gemini 3.1 Pro54.3
4Agent S3 + Qwen3-VL-235B-A22B46.7
5Agent S3 + Qwen3-VL-32B-Thinking37.4
6Agent S3 + Qwen3-VL-8B-Thinking28.6
7Agent S3 + InternVL3.5-14B13.5
8Agent S3 + InternVL3.5-8B9.6

Interactive version: theaggregate.ai/benchmark?slug=macagentbench-agent-s3 · How It Works · Data refreshed daily, snapshot 2026-09-26.