VibeSearchBench (ReAct) - Daily: leaderboard

Metric: Triplet F1 (%; an LLM judge matches the knowledge graph the agent extracts after a multi-turn session with a progressive-disclosure user simulator against the annotated ground-truth graph; ReAct loop with self-summarising context compaction; 100 everyday VibeSearch-Daily tasks; mean of 3 runs per task). Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1Claude Opus 4.625.95
2DeepSeek V4 Pro25.37
3Kimi K2.625.05
4Gemini 3.1 Pro (Preview)24.66
5Qwen 3.5 397B A17B22.02
6GPT-5.421.13
7Seed 2.0 Pro20.58

Interactive version: theaggregate.ai/benchmark?slug=vibesearchbench-react-daily · How It Works · Data refreshed daily, snapshot 2026-09-26.