CoLT-Drive: leaderboard

Metric: Pair accuracy (%; share of the 3,536 counterfactual samples, 1,768 scenes in both v_full with surrounding traffic and v_clean with background vehicles removed, where the model outputs an acceptable longitudinal-lateral meta-action pair for a front-view driving image with a counterfactually inserted rare object, ego-motion history and navigation command; human-adjudicated acceptable-pair sets; same structured prompt and greedy decoding for every model; decisions extracted by rule and normalized by a text-only DeepSeek-V4-Pro mapper, then scored deterministically; invalid outputs count as wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.583
2Qwen 3 VL 8B65.5

Interactive version: theaggregate.ai/benchmark?slug=colt-drive · How It Works · Data refreshed daily, snapshot 2026-09-26.