TimeLitmus - Traffic Hard Paired Contrasts: leaderboard

Metric: Hard paired contrast pair correctness (%; both endpoints correct on 60 naturally occurring similar Traffic pairs with different gold responses; ten LLMs with deterministic decoding where available, invalid or unparsable outputs scored incorrect). Source: arxiv.org. Saturation forecast: Around 2042. 10 models tracked.

Top models

#ModelScore
1Qwen 3.5 9B11.7
2DeepSeek V4 Flash11.7
3DeepSeek R1 052811.7
4Qwen 3.5 4B6.7
5MiniMax-M36.7
6Qwen Plus6.7
7Claude Sonnet 4.61.7
8GLM-51.7
9GPT-5.40
10Gemini 3.5 Flash0

Interactive version: theaggregate.ai/benchmark?slug=timelitmus-traffic-hard-paired-contrasts · How It Works · Data refreshed daily, snapshot 2026-09-26.