HippoCamp (Autonomous Agents) - Profiling: leaderboard

Metric: LLM-judge answer accuracy (%) on HippoCamp, the 60 user-profiling questions, over three personal file-system profiles (Bei, Adam, Victoria; 42.4 GB, about 2K files), each answer judged correct or not by GPT-4o against the reference answer; autonomous-agent setting: Qwen3-VL-8B-Instruct, Gemini 2.5 Flash and GPT-5.2 drive the benchmark vacuum-Docker terminal interface, and ChatGPT agent runs in its hosted product mode; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 4 models tracked.

Top models

#ModelScore
1GPT-5.230
2Gemini 2.5 Flash25
3Qwen 3 VL 8B Instruct16.7

Interactive version: theaggregate.ai/benchmark?slug=hippocamp-autonomous-agents-profiling · How It Works · Data refreshed daily, snapshot 2026-10-07.