HumanEval-Mul — leaderboard

A multilingual variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics.

Source: github.com.

Interactive version: theaggregate.ai/benchmark?slug=humaneval-mul · How the rankings work · Data refreshed daily, snapshot 2026-07-22.