AssertLLM2 - Proof Coverage: leaderboard
Metric: Proof coverage (%): share of the design logic the formal solver needs to prove the generated assertions, SystemVerilog assertions generated from the structured specification of each of 83 AssertLLM2 designs and checked in Cadence JasperGold, mean of three independent generations per design (the paper's default average setting); higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 23.3 |
| 2 | GPT-5.2 | 22.46 |
| 3 | Gemini 2.5 Pro | 20.19 |
| 4 | DeepSeek V3.2 | 16.75 |
| 5 | Qwen 3 Coder Plus | 12.77 |
| 6 | Llama 4 Maverick | 9.09 |
Interactive version: theaggregate.ai/benchmark?slug=assertllm2-proof-coverage · How It Works · Data refreshed daily, snapshot 2026-10-07.