DomainCodeBench: leaderboard
20 Python code-generation tasks, five each in healthcare (FHIR/HIPAA), finance (Black-Scholes, VaR), molecular simulation and legal documents; composite score with correctness weighted 40% (2026).
Metric: Composite Score (%). Source: huggingface.co. Status: saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 2.5 Coder 7B | 89.77 |
| 2 | starcoder2-15B | 88.96 |
Interactive version: theaggregate.ai/benchmark?slug=domaincodebench · How It Works · Data refreshed daily, snapshot 2026-09-05.