HORIZON - LLM Long-Horizon User Modeling: leaderboard

Metric: Recall@100 (%) of the user's next ten items when the LLM describes the ten items it expects the user to engage with, items retrieved from the Amazon Reviews catalog with the BLAIR item encoder and a FAISS index, zero-shot, on HORIZON's held-out set of one million out-of-distribution users after the 2020 cut-off; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 3 models tracked.

Top models

#ModelScore
1Qwen 3 8B15.75
2Llama 3.1 8B Instruct13.25
3Gemma 2 9B (IT)10.39

Interactive version: theaggregate.ai/benchmark?slug=horizon-llm-long-horizon-user-modeling · How It Works · Data refreshed daily, snapshot 2026-10-07.