AssertLLM2 - Formal Coverage: leaderboard
Metric: Formal RTL coverage (%) reached by the proven assertions, SystemVerilog assertions generated from the structured specification of each of 83 AssertLLM2 designs and checked in Cadence JasperGold, mean of three independent generations per design (the paper's default average setting); higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 22.38 |
| 2 | GPT-5.2 | 21.18 |
| 3 | Gemini 2.5 Pro | 19.16 |
| 4 | DeepSeek V3.2 | 15.73 |
| 5 | Qwen 3 Coder Plus | 12.07 |
| 6 | Llama 4 Maverick | 8.46 |
Interactive version: theaggregate.ai/benchmark?slug=assertllm2-formal-coverage · How It Works · Data refreshed daily, snapshot 2026-10-07.