ISA-Bench: leaderboard

Metric: Tasks solved (%; share of the 92 tasks of five programming games with constrained instruction sets (SIC-1, Human Resource Machine, TIS-100, MHRD, BOX-256) whose program passes every verification case in the game's own parser, virtual machine and verifier; up to five attempts with parse, runtime or per-test execution feedback; Ollama default builds, temperature 0.7, 4,096 output tokens). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.

Top models

#ModelScore
1ISA-Bench Gemma-4 31B (Thinking, Ollama build)60.87
2ISA-Bench Qwen3 32B (Thinking, Ollama build)31.52
3ISA-Bench Qwen3 14B (Thinking, Ollama build)26.09
4ISA-Bench Phi-4-Reasoning (Ollama build)26.09
5ISA-Bench DeepSeek-R1 32B (Ollama build)25
6ISA-Bench Devstral-Small-2 (Ollama build)23.91
7ISA-Bench Qwen3-Coder 30B (Ollama build)22.83
8ISA-Bench Qwen3 8B (Thinking, Ollama build)21.74
9ISA-Bench Qwen2.5-Coder 32B (Ollama build)14.13
10ISA-Bench Granite-4.1 8B (Ollama build)13.04

Interactive version: theaggregate.ai/benchmark?slug=isa-bench · How It Works · Data refreshed daily, snapshot 2026-09-26.