HumanEval-Mul — leaderboard
A multilingual variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics.
Source: github.com.
Interactive version: theaggregate.ai/benchmark?slug=humaneval-mul · How the rankings work · Data refreshed daily, snapshot 2026-07-22.