Interpretive Canons (BVerfG): leaderboard

Metric: Mean positive-class F1 over the seven binary subtasks (%; sentence-level classification of reading criteria, argument support and the four interpretive canons of Larenz in German Federal Constitutional Court decisions; expert-written prompt with worked examples (the same for every model), reasoning models at temperature 1.0; decision-disjoint held-out test split (60% of each subtask, 1:3 positives to negatives)). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro75.2
2MiniMax-M375.2
3DeepSeek V4 Flash73.5
4Gemini 3.5 Flash Lite72.7

Interactive version: theaggregate.ai/benchmark?slug=interpretive-canons-bverfg · How It Works · Data refreshed daily, snapshot 2026-09-26.