CrypFormBench - Interpretation: leaderboard

Metric: Interpretation score (0-100) on CrypFormBench (700 instances over 677 cryptographic schemes and 7 formal verifier languages: Scyther, Tamarin, AVISPA, ProVerif, Maude-NPA, CryptoVerif, EasyCrypt): explain a formal model in natural language (notation-level comments and a global scheme summary), scored 0.3 logic-description similarity + 0.3 annotation similarity (Qwen3-Embedding-8B cosine similarity to the verified references) + 0.4 the executability-based score of the annotated code, averaged over languages; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1GPT-4o94.1
2GPT-4o Mini93.8
3DeepSeek R177.4
4Gemini 2.5 Pro69.1
5GLM-462.3
6Grok 349.6

Interactive version: theaggregate.ai/benchmark?slug=crypformbench-interpretation · How It Works · Data refreshed daily, snapshot 2026-09-29.