MacAgentBench (OpenClaw): leaderboard

Metric: Pass@1 (%; share of the 676 macOS desktop tasks in 25 applications solved on a single attempt, judged by deterministic rule-based state checks; the OpenClaw in-container agent harness with shell, AppleScript and its skill library, at most 50 steps). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (OpenClaw)73.7
2Gemini 3.1 Pro (OpenClaw)63.3
3GPT-5.4 (OpenClaw)60.7
4Qwen3-VL-235B-A22B (OpenClaw)46
5Qwen3-VL-32B-Thinking (OpenClaw)33.7
6Qwen3-VL-8B-Thinking (OpenClaw)26.9
7InternVL3.5-14B (OpenClaw)0
8InternVL3.5-8B (OpenClaw)0

Interactive version: theaggregate.ai/benchmark?slug=macagentbench-openclaw · How It Works · Data refreshed daily, snapshot 2026-09-26.