MacAgentBench (Agent-S3) - System: leaderboard

Metric: Pass@1 (%; share of the 44 System tasks solved on a single attempt, deterministic rule-based checks; the Agent S3 multi-agent framework, at most 50 steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Agent S3 + Claude Opus 4.663.6
2Agent S3 + GPT-5.456.8
3Agent S3 + Qwen3-VL-32B-Thinking54.5
4Agent S3 + Gemini 3.1 Pro52.3
5Agent S3 + Qwen3-VL-235B-A22B38.6
6Agent S3 + Qwen3-VL-8B-Thinking25
7Agent S3 + InternVL3.5-14B15.9
8Agent S3 + InternVL3.5-8B9.1

Interactive version: theaggregate.ai/benchmark?slug=macagentbench-agent-s3-system · How It Works · Data refreshed daily, snapshot 2026-09-26.