MultiPL-E — leaderboard
MultiPL-E: Measures model capability on programming, code generation, code repair, or repository-level software tasks.
Source: arxiv.org.
Interactive version: theaggregate.ai/benchmark?slug=multipl-e · How the rankings work · Data refreshed daily, snapshot 2026-07-22.