TimeLitmus - Traffic Event-Side Counterfactual Pairs: leaderboard

Metric: Counterfactual pair correctness (%; both independently evaluated endpoints correct on 30 Traffic pairs whose incident severity or capacity impact is changed while the series is kept; ten LLMs with deterministic decoding where available, invalid or unparsable outputs scored incorrect). Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.

Top models

#ModelScore
1GPT-5.463.3
2MiniMax-M356.7
3GLM-556.7
4Claude Sonnet 4.650
5Qwen 3.5 4B50
6DeepSeek V4 Flash33.3
7DeepSeek R1 052826.7
8Qwen 3.5 9B23.3
9Qwen Plus13.3
10Gemini 3.5 Flash6.7

Interactive version: theaggregate.ai/benchmark?slug=timelitmus-traffic-event-side-counterfactual-pairs · How It Works · Data refreshed daily, snapshot 2026-09-26.