AsyncPlan Robo Challenge (Planner): leaderboard
Metric: Plan accuracy (%; share of the 140 Robo Challenge kitchen tasks (seven splits of 20 built on Robotouille, with object states, resources, stations and multiple agents) whose plan is valid and reaches the optimal makespan; the LLM writes the schedule directly (Planner); zero-shot, temperature 0). Source: arxiv.org. Saturation forecast: Around 2032. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.6 35B A3B | 59.3 |
| 2 | GPT-5 Mini | 46.4 |
| 3 | Gemini 3 Flash | 40.7 |
| 4 | DeepSeek V4 Flash | 12.9 |
Interactive version: theaggregate.ai/benchmark?slug=asyncplan-robo-challenge-planner · How It Works · Data refreshed daily, snapshot 2026-09-26.