BC-Bench - Test Generation (pass^5): leaderboard
Metric: pass^5 (%; share of tasks resolved in all five runs; test generation: a task is resolved when the new test fails on the base commit and passes with the gold patch; 101 manually curated Business Central (AL) tasks from two Microsoft production repositories; five independent runs per task, 30-minute timeout, agents see the AL code as plain text without the development environment; benchmark version as recorded per configuration (0.1.0 to 0.2.2)). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GitHub Copilot CLI + claude-opus-4.6 | 37.6 |
| 2 | GitHub Copilot CLI + claude-opus-4.5 | 20.8 |
| 3 | GitHub Copilot CLI + gpt-5.3-codex | 20.8 |
| 4 | GitHub Copilot CLI + gpt-5.2-codex | 16.8 |
Interactive version: theaggregate.ai/benchmark?slug=bc-bench-test-generation-pass-pow-5 · How It Works · Data refreshed daily, snapshot 2026-09-29.