DriveCombo - L4 Static and Dynamic Rules: leaderboard

Metric: Accuracy (%) on DriveCombo Level 4 (a static and a dynamic rule together) four-option multiple-choice questions on the 20% test split, rules from five countries' traffic codes; input is the question with four RGB frames of the simulated driving scene; zero-shot, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Pro69.82#75
2Claude Sonnet 4.569.62#138
3GLM-4.5V68.22#339
4Gemini 2.5 Pro68.03#145
5Gemini 2.5 Flash64.03#237
6GPT-5 Nano63.43#415
7Gemma 3 27B63.1#596
8GPT-5 Mini62.27#176
9Gemma 3 12B52.32#666
10Gemma 3 4B44.55#1084
11Llama 3.2 11B41.94#1183

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=drivecombo-l4-static-and-dynamic-rules · How It Works · Data refreshed daily, snapshot 2026-10-11.