MacAgentBench (Agent-S3): leaderboard
Metric: Pass@1 (%; share of the 676 macOS desktop tasks in 25 applications solved on a single attempt, judged by deterministic rule-based state checks; the Agent S3 multi-agent framework, at most 50 steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Agent S3 + Claude Opus 4.6 | 66.9 |
| 2 | Agent S3 + GPT-5.4 | 58.9 |
| 3 | Agent S3 + Gemini 3.1 Pro | 54.3 |
| 4 | Agent S3 + Qwen3-VL-235B-A22B | 46.7 |
| 5 | Agent S3 + Qwen3-VL-32B-Thinking | 37.4 |
| 6 | Agent S3 + Qwen3-VL-8B-Thinking | 28.6 |
| 7 | Agent S3 + InternVL3.5-14B | 13.5 |
| 8 | Agent S3 + InternVL3.5-8B | 9.6 |
Interactive version: theaggregate.ai/benchmark?slug=macagentbench-agent-s3 · How It Works · Data refreshed daily, snapshot 2026-09-26.