BC-Bench - Test Generation: leaderboard
Metric: Mean resolution rate (%; mean over five runs; test generation: a task is resolved when the new test fails on the base commit and passes with the gold patch; 101 manually curated Business Central (AL) tasks from two Microsoft production repositories; five independent runs per task, 30-minute timeout, agents see the AL code as plain text without the development environment; benchmark version as recorded per configuration (0.1.0 to 0.2.2)). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GitHub Copilot CLI + claude-opus-4.6 | 60.4 |
| 2 | GitHub Copilot CLI + claude-opus-4.5 | 45.5 |
| 3 | GitHub Copilot CLI + gpt-5.3-codex | 45.3 |
| 4 | GitHub Copilot CLI + gpt-5.2-codex | 44 |
Interactive version: theaggregate.ai/benchmark?slug=bc-bench-test-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.