TRACK - Code Generation (Facts in Context): leaderboard

Metric: Holistic pass (%; BigCodeBench-derived coding problems that require specific library API versions, 500 items; the final answer is correct and the reasoning entails every required fact; answered with the model's own missing facts appended to the prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini36.8
2Qwen 3 4B28.2
3Qwen 3 8B27.8
4Qwen 3 1.7B18.3

Interactive version: theaggregate.ai/benchmark?slug=track-code-generation-facts-in-context · How It Works · Data refreshed daily, snapshot 2026-09-26.