Carlini Applied LLM Benchmark — leaderboard

Nicholas Carlini's practical coding benchmark with ~100 tests covering real-world tasks like parsing BNF grammars, writing shell one-liners, and generating C programs.

Metric: Pass Rate (%). Source: github.com. Status: saturation imminent. 8 models tracked.

Top models

#ModelScore
1O1 Mini62
2Claude 3.5 Sonnet56
3GPT-4o48
4Gemini 1.5 Pro43
5Claude 3 Opus42
6GPT-4o Mini36
7Mistral Large28
8GPT-3.526

Interactive version: theaggregate.ai/benchmark?slug=carlini-applied-llm-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.