CrypFormBench - Interpretation: leaderboard
Metric: Interpretation score (0-100) on CrypFormBench (700 instances over 677 cryptographic schemes and 7 formal verifier languages: Scyther, Tamarin, AVISPA, ProVerif, Maude-NPA, CryptoVerif, EasyCrypt): explain a formal model in natural language (notation-level comments and a global scheme summary), scored 0.3 logic-description similarity + 0.3 annotation similarity (Qwen3-Embedding-8B cosine similarity to the verified references) + 0.4 the executability-based score of the annotated code, averaged over languages; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 94.1 |
| 2 | GPT-4o Mini | 93.8 |
| 3 | DeepSeek R1 | 77.4 |
| 4 | Gemini 2.5 Pro | 69.1 |
| 5 | GLM-4 | 62.3 |
| 6 | Grok 3 | 49.6 |
Interactive version: theaggregate.ai/benchmark?slug=crypformbench-interpretation · How It Works · Data refreshed daily, snapshot 2026-09-29.