Carlini Applied LLM Benchmark — leaderboard
Nicholas Carlini's practical coding benchmark with ~100 tests covering real-world tasks like parsing BNF grammars, writing shell one-liners, and generating C programs.
Metric: Pass Rate (%). Source: github.com. Status: saturation imminent. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O1 Mini | 62 |
| 2 | Claude 3.5 Sonnet | 56 |
| 3 | GPT-4o | 48 |
| 4 | Gemini 1.5 Pro | 43 |
| 5 | Claude 3 Opus | 42 |
| 6 | GPT-4o Mini | 36 |
| 7 | Mistral Large | 28 |
| 8 | GPT-3.5 | 26 |
Interactive version: theaggregate.ai/benchmark?slug=carlini-applied-llm-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.