CrypFormBench - Generation: leaderboard
Metric: Generation score (0-100) on CrypFormBench (700 instances over 677 cryptographic schemes and 7 formal verifier languages: Scyther, Tamarin, AVISPA, ProVerif, Maude-NPA, CryptoVerif, EasyCrypt): write a tool-specific formal model of a described scheme; the harmonic mean of the tool-executable rate and the harmonic mean of verification-outcome accuracy and F1 against the gold tool verdicts, averaged over languages; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 3 | 12.8 |
| 2 | GPT-4o | 9.5 |
| 3 | Gemini 2.5 Pro | 7.7 |
| 4 | DeepSeek R1 | 5.8 |
| 5 | GPT-4o Mini | 0 |
| 6 | GLM-4 | 0 |
Interactive version: theaggregate.ai/benchmark?slug=crypformbench-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.