PinchBench — leaderboard

Open-source Kilo.ai benchmark where coding agents complete practical automated tasks such as creating files, researching data, and producing working app artifacts.

Metric: Success Rate (%). Source: pinchbench.com. Status: saturated. 58 models tracked.

Top models

#ModelScore
1Claude Opus 4.6100
2Mercury 2100
3Qwen 3.7 Max93.44
4Grok Build 0.192.07
5MiMo-V2.591.87
6Claude Opus 4.891.76
7Claude Opus 4.791.58
8DeepSeek V4 Flash91.46
9Grok 4.591.17
10GPT-5.6 Luna90.79
11Nemotron 3 Ultra90.58
12Claude Haiku 4.590.4
13Seed 2.0 Lite89.71
14MiMo-V2.5-Pro89.52
15GPT-5.588.96

Interactive version: theaggregate.ai/benchmark?slug=pinchbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.