PersonalHomeBench - Information Grounding (Sole Reasoning): leaderboard

Metric: Accuracy (%) on 1,000 information grounding questions (factual lookup to multi-hop reasoning over household entities; exact match, else a Gemini 2.5 Pro semantic-equivalence judge), over the text-based split of PersonalHomeBench (1,100 generated, manually reviewed households with persona-based occupants and appliances), temperature 0, in the sole-reasoning setting (no tools; the full household context, occupant profiles and memories are given in the prompt); each model at its stated reasoning mode; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GPT-OSS-20B (High)81.2
2GPT-4o80
3GPT-OSS-20B (Low)80
4GPT-OSS-20B79.94
5GPT-477.9
6Qwen 3 4B (Reasoning)75.7
7nemotron-3-nano-30B-a3B75.62
8Qwen 3 4B (Non-reasoning)75

Interactive version: theaggregate.ai/benchmark?slug=personalhomebench-information-grounding-sole-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.