BC-Bench - Test Generation (pass^5): leaderboard

Metric: pass^5 (%; share of tasks resolved in all five runs; test generation: a task is resolved when the new test fails on the base commit and passes with the gold patch; 101 manually curated Business Central (AL) tasks from two Microsoft production repositories; five independent runs per task, 30-minute timeout, agents see the AL code as plain text without the development environment; benchmark version as recorded per configuration (0.1.0 to 0.2.2)). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1GitHub Copilot CLI + claude-opus-4.637.6
2GitHub Copilot CLI + claude-opus-4.520.8
3GitHub Copilot CLI + gpt-5.3-codex20.8
4GitHub Copilot CLI + gpt-5.2-codex16.8

Interactive version: theaggregate.ai/benchmark?slug=bc-bench-test-generation-pass-pow-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.