SRBench - ML-100K (Recall@5): leaderboard

Metric: Recall@5 (fraction) of the held-out next item among five recommendations on ML-100K, LLM sequential recommendation with the full-length interaction history and SRBench's augmented prompt and extractor; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 8 models tracked.

Top models

#ModelScore
1Claude Sonnet 40.03
2GPT-4.10.03
3Claude Sonnet 4 (Thinking)0.03
4Grok 30.03

Interactive version: theaggregate.ai/benchmark?slug=srbench-ml-100k-recall-5 · How It Works · Data refreshed daily, snapshot 2026-10-07.