DriveCombo-Text - L4 Static and Dynamic Rules: leaderboard

Metric: Accuracy (%) on DriveCombo Level 4 (a static and a dynamic rule together) four-option multiple-choice questions on the 20% test split, rules from five countries' traffic codes; text-only variant in which the frames are replaced by a textual scene description; zero-shot, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Pro75.78#75
2Claude Sonnet 4.574.38#138
3Gemini 2.5 Pro73.76#145
4GLM-4.5V71.94#339
5Gemini 2.5 Flash69.33#237
6Gemma 3 27B67.55#596
7GPT-5 Mini66.95#176
8GPT-5 Nano66.82#415
9Gemma 3 12B56.56#666
10Gemma 3 4B48.16#1084
11Llama 3.2 11B45.59#1183

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=drivecombo-text-l4-static-and-dynamic-rules · How It Works · Data refreshed daily, snapshot 2026-10-11.