HippoCamp (Autonomous Agents): leaderboard
Metric: LLM-judge answer accuracy (%) on HippoCamp, all 581 queries (60 profiling and 521 factual-retention questions, sample-weighted), over three personal file-system profiles (Bei, Adam, Victoria; 42.4 GB, about 2K files), each answer judged correct or not by GPT-4o against the reference answer; autonomous-agent setting: Qwen3-VL-8B-Instruct, Gemini 2.5 Flash and GPT-5.2 drive the benchmark vacuum-Docker terminal interface, and ChatGPT agent runs in its hosted product mode; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 44.1 |
| 2 | Gemini 2.5 Flash | 22 |
| 3 | Qwen 3 VL 8B Instruct | 11.7 |
Interactive version: theaggregate.ai/benchmark?slug=hippocamp-autonomous-agents · How It Works · Data refreshed daily, snapshot 2026-10-07.