UniClawBench - Cross-Platform: leaderboard

Metric: Pass rate (%) on the 80 cross-platform tasks of 400 bilingual real-world tasks run live in Docker containers under the OpenClaw framework (v2026.3.11), with up to two simulated user follow-ups and a hidden GPT-5.4 Codex supervisor that grades step checkpoints. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.838.8
2GPT-5.435
3Claude Sonnet 4.633.8
4Kimi K2.623.7
5Qwen 3.5 Plus22.5
6Gemini 3.1 Pro (Preview)12.5
7Gemini 3 Flash7.5
8GPT-5.4 Mini3.7
9GPT-4.12.5
10Gemini 3.1 Flash Lite2.5

Interactive version: theaggregate.ai/benchmark?slug=uniclawbench-cross-platform · How It Works · Data refreshed daily, snapshot 2026-09-29.