CoLT-Drive - False Positive: leaderboard
Metric: Pair accuracy (%; the 435 v_full samples, naturalistic surrounding traffic kept, whose inserted object is one of the benign lookalikes that should not change the action (shadows, puddles, flat cardboard, leaves), where the model outputs an acceptable longitudinal-lateral meta-action pair for a front-view driving image with a counterfactually inserted rare object, ego-motion history and navigation command; human-adjudicated acceptable-pair sets; same structured prompt and greedy decoding for every model; decisions extracted by rule and normalized by a text-only DeepSeek-V4-Pro mapper, then scored deterministically; invalid outputs count as wrong). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 91.5 |
| 2 | Qwen 3 VL 8B | 84.4 |
Interactive version: theaggregate.ai/benchmark?slug=colt-drive-false-positive · How It Works · Data refreshed daily, snapshot 2026-09-26.