VibeMemBench (Frozen Verified Experience): leaderboard

Metric: Resolved (%; the frozen verified experience for each target, distilled from earlier trajectories in the same repository and verified to help a reference solver, injected into the solver context; share of 444 target-seed runs (111 SWE-style repository coding targets from 90 SWE-rebench V2 repositories x 4 seeds) whose declared tests pass; MiniSWEAgent solver). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1GLM-5.2 (Mini-SWE-Agent)80.4
2Qwen3.8-Max (Mini-SWE-Agent)80.2
3Kimi K2.7 Code (Mini-SWE-Agent)74.3
4DeepSeek-V4-Pro (Mini-SWE-Agent)67.1
5GLM-5 (Mini-SWE-Agent)57.4

Interactive version: theaggregate.ai/benchmark?slug=vibemembench-frozen-verified-experience · How It Works · Data refreshed daily, snapshot 2026-09-26.