PRISM-BN - CPD-KL: leaderboard

Metric: CPD KL divergence (column-wise KL from the reference conditional probability table to the predicted one, Laplace-smoothed with epsilon 1e-6 and averaged over parent-state columns, then over all correctly recovered edges of all five domains; lower is better; zero-shot extraction of a parameterized Bayesian network from each of the 5,054 PRISM-BN text descriptions through the four-phase prompt pipeline; predicted node and state names are aligned one-to-one with the reference by a Llama 3.3 70B semantic judge). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Claude Haiku 4.51.11
2Llama 4 Maverick1.14
3DeepSeek V31.52
4Qwen 3 30B A3B 2507 Instruct1.72
5GPT-4o Mini2.25
6Gemma 3 12B (IT)3.14

Interactive version: theaggregate.ai/benchmark?slug=prism-bn-cpd-kl · How It Works · Data refreshed daily, snapshot 2026-09-26.