DriveCombo - L2 Static Rule Integration: leaderboard

Metric: Accuracy (%) on DriveCombo Level 2 (two coexisting static rules, such as signs or speed limits) four-option multiple-choice questions on the 20% test split, rules from five countries' traffic codes; input is the question with four RGB frames of the simulated driving scene; zero-shot, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Pro83.66#75
2Claude Sonnet 4.582.16#138
3GLM-4.5V78.49#339
4Gemini 2.5 Pro77.19#145
5GPT-5 Mini72.59#176
6Gemini 2.5 Flash72.07#237
7GPT-5 Nano69.76#415
8Gemma 3 27B67.39#596
9Gemma 3 12B58.69#666
10Llama 3.2 11B49.23#1183
11Gemma 3 4B47.8#1084

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=drivecombo-l2-static-rule-integration · How It Works · Data refreshed daily, snapshot 2026-10-11.