AsyncPlan Online Robo Challenge (CP-SAT Formalizer): leaderboard
Metric: Plan accuracy (%; share of the 140 Online Robo Challenge episodes of the kitchen tasks (seven splits of 20 built on Robotouille, with object states, resources, stations and multiple agents) whose plan is valid and reaches the optimal makespan; execution-time events (new deliveries, deadlines, resource changes) require replanning or one-shot re-formalization; the LLM writes a CP-SAT constraint program that OR-Tools solves (CP-SAT Formalizer); zero-shot, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 83.6 |
| 2 | Qwen 3.6 35B A3B | 72.1 |
| 3 | GPT-5 Mini | 25 |
| 4 | DeepSeek V4 Flash | 3.6 |
Interactive version: theaggregate.ai/benchmark?slug=asyncplan-online-robo-challenge-cp-sat-formalizer · How It Works · Data refreshed daily, snapshot 2026-09-26.