BC-Bench - Test Generation: leaderboard

Metric: Mean resolution rate (%; mean over five runs; test generation: a task is resolved when the new test fails on the base commit and passes with the gold patch; 101 manually curated Business Central (AL) tasks from two Microsoft production repositories; five independent runs per task, 30-minute timeout, agents see the AL code as plain text without the development environment; benchmark version as recorded per configuration (0.1.0 to 0.2.2)). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1GitHub Copilot CLI + claude-opus-4.660.4
2GitHub Copilot CLI + claude-opus-4.545.5
3GitHub Copilot CLI + gpt-5.3-codex45.3
4GitHub Copilot CLI + gpt-5.2-codex44

Interactive version: theaggregate.ai/benchmark?slug=bc-bench-test-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.