HippoCamp (Search Agents) - Profiling: leaderboard
Metric: LLM-judge answer accuracy (%) on HippoCamp, the 60 user-profiling questions, over three personal file-system profiles (Bei, Adam, Victoria; 42.4 GB, about 2K files), each answer judged correct or not by GPT-4o against the reference answer; search-agent setting: each model runs a search loop against the profile-local retrieval server (ReAct prompting for Gemini 2.5 Flash and Qwen3-30B-A3B, the trained Search-R1 loop for Search-R1); higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 20 |
| 2 | Qwen 3 30B A3B | 13.5 |
Interactive version: theaggregate.ai/benchmark?slug=hippocamp-search-agents-profiling · How It Works · Data refreshed daily, snapshot 2026-10-07.