PRISM-BN - State F1: leaderboard
Metric: State F1 (%; macro-averaged over nodes matched in the node phase, so conditional on node alignment; zero-shot extraction of a parameterized Bayesian network from each of the 5,054 PRISM-BN text descriptions through the four-phase prompt pipeline; predicted node and state names are aligned one-to-one with the reference by a Llama 3.3 70B semantic judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 93.1 |
| 2 | DeepSeek V3 | 91.8 |
| 3 | Qwen 3 30B A3B 2507 Instruct | 90.7 |
| 4 | GPT-4o Mini | 90.3 |
| 5 | Gemma 3 12B (IT) | 89.3 |
| 6 | Llama 4 Maverick | 88.6 |
Interactive version: theaggregate.ai/benchmark?slug=prism-bn-state-f1 · How It Works · Data refreshed daily, snapshot 2026-09-26.