ISA-Bench - MHRD: leaderboard
Metric: Tasks solved (%; share of the 8 MHRD tasks (composing circuits from NAND gates in a hardware description language, checked against complete truth tables) whose program passes every verification case in the game's own parser, virtual machine and verifier; up to five attempts with parse, runtime or per-test execution feedback; Ollama default builds, temperature 0.7, 4,096 output tokens). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | ISA-Bench Gemma-4 31B (Thinking, Ollama build) | 100 |
| 2 | ISA-Bench Qwen3 32B (Thinking, Ollama build) | 87.5 |
| 3 | ISA-Bench Qwen3 14B (Thinking, Ollama build) | 87.5 |
| 4 | ISA-Bench DeepSeek-R1 32B (Ollama build) | 87.5 |
| 5 | ISA-Bench Devstral-Small-2 (Ollama build) | 75 |
| 6 | ISA-Bench Qwen3-Coder 30B (Ollama build) | 62.5 |
| 7 | ISA-Bench Qwen2.5-Coder 32B (Ollama build) | 62.5 |
| 8 | ISA-Bench Granite-4.1 8B (Ollama build) | 62.5 |
| 9 | ISA-Bench Qwen3 8B (Thinking, Ollama build) | 50 |
| 10 | ISA-Bench Phi-4 14B (Ollama build) | 37.5 |
Interactive version: theaggregate.ai/benchmark?slug=isa-bench-mhrd · How It Works · Data refreshed daily, snapshot 2026-09-26.