EsoLang-Bench — leaderboard
Tests genuine reasoning via code generation in 5 esoteric programming languages (Brainfuck, Befunge-98, Whitespace, Unlambda, Shakespeare). 80 problems per language with 6 test cases each. Frontier models score ~90% on equivalent Python tasks but only ~4% here.
Metric: Accuracy (%). Source: esolang-bench.vercel.app. Status: saturation imminent. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 4.2 |
| 2 | O4 Mini | 3.2 |
| 3 | Gemini 3 Pro | 2.7 |
| 4 | Qwen 3 235B A22B | 1 |
| 5 | Kimi K2 | 0.7 |
Interactive version: theaggregate.ai/benchmark?slug=esolang-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.