HippoCamp (Search Agents): leaderboard

Metric: LLM-judge answer accuracy (%) on HippoCamp, all 581 queries (60 profiling and 521 factual-retention questions, sample-weighted), over three personal file-system profiles (Bei, Adam, Victoria; 42.4 GB, about 2K files), each answer judged correct or not by GPT-4o against the reference answer; search-agent setting: each model runs a search loop against the profile-local retrieval server (ReAct prompting for Gemini 2.5 Flash and Qwen3-30B-A3B, the trained Search-R1 loop for Search-R1); higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 3 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash33.6
2Qwen 3 30B A3B25.3

Interactive version: theaggregate.ai/benchmark?slug=hippocamp-search-agents · How It Works · Data refreshed daily, snapshot 2026-10-07.