PersonalHomeBench - Feature Recommendation (Agentic): leaderboard

Metric: MAP@1 (times 100) on 1,000 feature recommendation tasks (rank appliance features from API-level inventories by contextual relevance), over the text-based split of PersonalHomeBench (1,100 generated, manually reviewed households with persona-based occupants and appliances), temperature 0, in the agentic setting (the model must gather household information and act through the benchmark's PersonalHomeTools toolbox, up to 15 turns); each model at its stated reasoning mode; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 11 models tracked.

Top models

#ModelScore
1GPT-424.7
2nemotron-3-nano-30B-a3B23.9
3GPT-OSS-20B (High)22.9
4Qwen 3 4B (Reasoning)22.9
5GPT-4o22.7
6GPT-OSS-20B21.7
7Qwen 3 4B (Non-reasoning)21.6
8GPT-OSS-20B (Low)21.1

Interactive version: theaggregate.ai/benchmark?slug=personalhomebench-feature-recommendation-agentic · How It Works · Data refreshed daily, snapshot 2026-10-07.