DriveCombo - L1 Atomic Rules: leaderboard

Metric: Accuracy (%) on DriveCombo Level 1 (one atomic traffic rule) four-option multiple-choice questions on the 20% test split, rules from five countries' traffic codes; input is the question with four RGB frames of the simulated driving scene; zero-shot, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Pro86.91#75
2Gemini 2.5 Pro85.71#145
3Claude Sonnet 4.583.8#138
4GPT-5 Mini80.62#176
5GLM-4.5V80.44#339
6Gemini 2.5 Flash79.31#237
7GPT-5 Nano74.04#415
8Gemma 3 27B73.94#596
9Gemma 3 12B64.32#666
10Gemma 3 4B58.79#1084
11Llama 3.2 11B52.46#1183

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=drivecombo-l1-atomic-rules · How It Works · Data refreshed daily, snapshot 2026-10-11.