PRISM-BN - Edge F1: leaderboard

Metric: Edge F1 (%; directed edges whose endpoints were both matched, no credit for reversed edges, so conditional on node alignment; zero-shot extraction of a parameterized Bayesian network from each of the 5,054 PRISM-BN text descriptions through the four-phase prompt pipeline; predicted node and state names are aligned one-to-one with the reference by a Llama 3.3 70B semantic judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1DeepSeek V396.6
2Llama 4 Maverick96
3Claude Haiku 4.595.9
4Qwen 3 30B A3B 2507 Instruct93.9
5Gemma 3 12B (IT)93.2
6GPT-4o Mini90.2

Interactive version: theaggregate.ai/benchmark?slug=prism-bn-edge-f1 · How It Works · Data refreshed daily, snapshot 2026-09-26.