AsyncPlan Online Robo Challenge (CP-SAT Formalizer): leaderboard

Metric: Plan accuracy (%; share of the 140 Online Robo Challenge episodes of the kitchen tasks (seven splits of 20 built on Robotouille, with object states, resources, stations and multiple agents) whose plan is valid and reaches the optimal makespan; execution-time events (new deliveries, deadlines, resource changes) require replanning or one-shot re-formalization; the LLM writes a CP-SAT constraint program that OR-Tools solves (CP-SAT Formalizer); zero-shot, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1Gemini 3 Flash83.6
2Qwen 3.6 35B A3B72.1
3GPT-5 Mini25
4DeepSeek V4 Flash3.6

Interactive version: theaggregate.ai/benchmark?slug=asyncplan-online-robo-challenge-cp-sat-formalizer · How It Works · Data refreshed daily, snapshot 2026-09-26.