AsyncPlan Online Robo Challenge (Planner): leaderboard
Metric: Plan accuracy (%; share of the 140 Online Robo Challenge episodes of the kitchen tasks (seven splits of 20 built on Robotouille, with object states, resources, stations and multiple agents) whose plan is valid and reaches the optimal makespan; execution-time events (new deliveries, deadlines, resource changes) require replanning or one-shot re-formalization; the LLM writes the schedule directly (Planner); zero-shot, temperature 0). Source: arxiv.org. Saturation forecast: Around 2035. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 Mini | 35 |
| 2 | Qwen 3.6 35B A3B | 30.7 |
| 3 | Gemini 3 Flash | 17.1 |
| 4 | DeepSeek V4 Flash | 12.9 |
Interactive version: theaggregate.ai/benchmark?slug=asyncplan-online-robo-challenge-planner · How It Works · Data refreshed daily, snapshot 2026-09-26.