DriveCombo-Text - L5 Rule Conflicts: leaderboard

Metric: Accuracy (%) on DriveCombo Level 5 (two conflicting rules resolved by the legal priority order) four-option multiple-choice questions on the 20% test split, rules from five countries' traffic codes; text-only variant in which the frames are replaced by a textual scene description; zero-shot, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Pro47.42#75
2Claude Sonnet 4.546.82#138
3Gemini 2.5 Pro46.6#145
4GLM-4.5V45.49#339
5Gemini 2.5 Flash42.77#237
6GPT-5 Mini41.47#176
7Gemma 3 27B41.15#596
8GPT-5 Nano36.33#415
9Gemma 3 12B33.08#666
10Llama 3.2 11B30.22#1183
11Gemma 3 4B29.82#1084

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=drivecombo-text-l5-rule-conflicts · How It Works · Data refreshed daily, snapshot 2026-10-11.