EsoLang-Bench — leaderboard

Tests genuine reasoning via code generation in 5 esoteric programming languages (Brainfuck, Befunge-98, Whitespace, Unlambda, Shakespeare). 80 problems per language with 6 test cases each. Frontier models score ~90% on equivalent Python tasks but only ~4% here.

Metric: Accuracy (%). Source: esolang-bench.vercel.app. Status: saturation imminent. 5 models tracked.

Top models

#ModelScore
1GPT-5.24.2
2O4 Mini3.2
3Gemini 3 Pro2.7
4Qwen 3 235B A22B1
5Kimi K20.7

Interactive version: theaggregate.ai/benchmark?slug=esolang-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.