RecRM-Bench - Instruction Following: leaderboard

Metric: Accuracy (%) of the overall instruction-following compliance score the model assigns to a recommender response under the expert rubrics (role, format, process, constraint, content quality, style), zero-shot, the model acting as a reward model for an agentic recommender on RecRM-Bench (real Meituan query-response logs), its parsed score compared with the gold label; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 7 models tracked.

Top models

#ModelScore
1LongCat-Flash-Chat64.84
2GPT-4.155.47
3LongCat Flash (Thinking)43.75
4DeepSeek V3.2 (Thinking)35.29
5DeepSeek V3.2 (Non-reasoning)30.47
6Qwen 3 Max (Thinking)26.67

Interactive version: theaggregate.ai/benchmark?slug=recrm-bench-instruction-following · How It Works · Data refreshed daily, snapshot 2026-10-07.