MacAgentBench (Agent-S3) - Development: leaderboard

Metric: Pass@1 (%; share of the 44 Development tasks solved on a single attempt, deterministic rule-based checks; the Agent S3 multi-agent framework, at most 50 steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Agent S3 + Claude Opus 4.688.6
2Agent S3 + GPT-5.475
3Agent S3 + Gemini 3.1 Pro72.7
4Agent S3 + Qwen3-VL-235B-A22B40.9
5Agent S3 + Qwen3-VL-32B-Thinking34.1
6Agent S3 + Qwen3-VL-8B-Thinking25
7Agent S3 + InternVL3.5-14B11.4
8Agent S3 + InternVL3.5-8B4.5

Interactive version: theaggregate.ai/benchmark?slug=macagentbench-agent-s3-development · How It Works · Data refreshed daily, snapshot 2026-09-26.