HippoCamp (Search Agents) - Factual Retention: leaderboard
Metric: LLM-judge answer accuracy (%) on HippoCamp, the 521 factual-retention questions (sample-weighted), over three personal file-system profiles (Bei, Adam, Victoria; 42.4 GB, about 2K files), each answer judged correct or not by GPT-4o against the reference answer; search-agent setting: each model runs a search loop against the profile-local retrieval server (ReAct prompting for Gemini 2.5 Flash and Qwen3-30B-A3B, the trained Search-R1 loop for Search-R1); higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 35.1 |
| 2 | Qwen 3 30B A3B | 26.7 |
Interactive version: theaggregate.ai/benchmark?slug=hippocamp-search-agents-factual-retention · How It Works · Data refreshed daily, snapshot 2026-10-07.