HippoCamp (Autonomous Agents) - Profiling: leaderboard
Metric: LLM-judge answer accuracy (%) on HippoCamp, the 60 user-profiling questions, over three personal file-system profiles (Bei, Adam, Victoria; 42.4 GB, about 2K files), each answer judged correct or not by GPT-4o against the reference answer; autonomous-agent setting: Qwen3-VL-8B-Instruct, Gemini 2.5 Flash and GPT-5.2 drive the benchmark vacuum-Docker terminal interface, and ChatGPT agent runs in its hosted product mode; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 30 |
| 2 | Gemini 2.5 Flash | 25 |
| 3 | Qwen 3 VL 8B Instruct | 16.7 |
Interactive version: theaggregate.ai/benchmark?slug=hippocamp-autonomous-agents-profiling · How It Works · Data refreshed daily, snapshot 2026-10-07.