TuRTLe Code Completion (Icarus Verilog) — leaderboard

TuRTLe leaderboard variant for RTL code completion evaluated with Icarus Verilog across VerilogEval MC and VeriGen.

Metric: Aggregated Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 44 models tracked.

Top models

#ModelScore
1GLM-583.98
2Gemma 4 31B82.57
3DeepSeek R1 052878.86
4GPT-OSS-120B77.82
5DeepSeek V3.1 Terminus76.57
6Qwen 3 235B A22B67.54
7GPT-OSS-20B66.48
8Llama 3.1 405B Instruct54.74
9Qwen 2.5 72B Instruct50.41
10Qwen 3 8B48.16
11Qwen 2.5 32B41.4
12QwQ-32B40.05
13starchat2-15B-v0.138.19
14DeepSeek R1 Distill Qwen 14B24.57
15Magistral Small22.62

Interactive version: theaggregate.ai/benchmark?slug=turtle-code-completion-icarus-verilog · How the rankings work · Data refreshed daily, snapshot 2026-07-22.