CDR-Bench - Order-Sensitive Filter Consistency: leaderboard

Metric: Group-level order consistency OCS@3 (%) on the filter-order-sensitive recipe groups of CDR-Bench: a group counts only when the model gets every permutation of the same operators right (filter placed before, between and after the mappers) within 3 attempts; non-thinking mode, baseline prompt; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 11 models tracked.

Top models

#ModelScore
1Claude Opus 4.618.68
2GPT-5.4 (Non-reasoning)9.46
3Qwen 3.6 Max8.27
4Claude Opus 4.58.14
5Qwen 3.6 27B (Non-reasoning)7.09
6GLM-5 (Non-reasoning)6.47
7Gemma 4 31B (IT) (Non-reasoning)6.11
8Kimi K2.6 (Non-reasoning)5.88
9Qwen 3.6 35B A3B (Non-reasoning)3.7
10DeepSeek V4 Flash (Non-reasoning)2.4
11Llama 4 Scout0.86

Interactive version: theaggregate.ai/benchmark?slug=cdr-bench-order-sensitive-filter-consistency · How It Works · Data refreshed daily, snapshot 2026-09-29.