Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

An Empirical Study of Proactive Coding Assista — leaderboard

Metric: Pass@1 (self-reported). Source: benchmarklist.com. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.613.57
2Gemini 3.1 Pro (Preview)11.46
3Qwen 3.5 397B A17B8.43
4DeepSeek V3.27.77
5GLM-57.77
6GPT-5.46.59
7MiniMax-M2.52.77

Interactive version: theaggregate.ai/benchmark?slug=an-empirical-study-of-proactive-coding-assista · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.