BackendForge: leaderboard

Metric: Task success rate (%; final oracle of 7,890 items including 640 reviewed co-evolved items; 56 contract-defined backend services rewritten from open-source applications; from a visible spec and OpenAPI contract the agent builds a Dockerized Python service, scored only through black-box HTTP pytest items; a task counts only when every item passes; one attempt per task in the same mini-SWE-style coding harness). Source: arxiv.org. Saturation forecast: Around August 2027. 14 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)28.57
2Claude Opus 4.7 (Max)17.86
3Gemini 3.5 Flash5.36
4GLM-5.13.57
5DeepSeek V4 Pro (Max)3.57
6Claude Sonnet 4.6 (Max)3.57
7Kimi K2.61.79
8DeepSeek V4 Pro (High)1.79
9MiniMax-M2.70
10DeepSeek V4 Flash (Max)0
11DeepSeek V4 Flash (High)0
12Hy3-preview (High)0

Interactive version: theaggregate.ai/benchmark?slug=backendforge · How It Works · Data refreshed daily, snapshot 2026-09-29.