DriveCombo-Text - L3 Dynamic Interactions: leaderboard

Metric: Accuracy (%) on DriveCombo Level 3 (two coexisting rules about other road participants) four-option multiple-choice questions on the 20% test split, rules from five countries' traffic codes; text-only variant in which the frames are replaced by a textual scene description; zero-shot, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.577.9#138
2GPT-5 Pro77.53#75
3Gemini 2.5 Pro75.41#145
4GLM-4.5V75.07#339
5GPT-5 Mini73.44#176
6GPT-5 Nano72.42#415
7Gemini 2.5 Flash72.04#237
8Gemma 3 27B69.51#596
9Gemma 3 12B59.39#666
10Llama 3.2 11B51.36#1183
11Gemma 3 4B50.93#1084

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=drivecombo-text-l3-dynamic-interactions · How It Works · Data refreshed daily, snapshot 2026-10-11.