DomainCodeBench: leaderboard

20 Python code-generation tasks, five each in healthcare (FHIR/HIPAA), finance (Black-Scholes, VaR), molecular simulation and legal documents; composite score with correctness weighted 40% (2026).

Metric: Composite Score (%). Source: huggingface.co. Status: saturated. 4 models tracked.

Top models

#ModelScore
1Qwen 2.5 Coder 7B89.77
2starcoder2-15B88.96

Interactive version: theaggregate.ai/benchmark?slug=domaincodebench · How It Works · Data refreshed daily, snapshot 2026-09-05.