RecRM-Bench - Instruction Following: leaderboard
Metric: Accuracy (%) of the overall instruction-following compliance score the model assigns to a recommender response under the expert rubrics (role, format, process, constraint, content quality, style), zero-shot, the model acting as a reward model for an agentic recommender on RecRM-Bench (real Meituan query-response logs), its parsed score compared with the gold label; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | LongCat-Flash-Chat | 64.84 |
| 2 | GPT-4.1 | 55.47 |
| 3 | LongCat Flash (Thinking) | 43.75 |
| 4 | DeepSeek V3.2 (Thinking) | 35.29 |
| 5 | DeepSeek V3.2 (Non-reasoning) | 30.47 |
| 6 | Qwen 3 Max (Thinking) | 26.67 |
Interactive version: theaggregate.ai/benchmark?slug=recrm-bench-instruction-following · How It Works · Data refreshed daily, snapshot 2026-10-07.