ActiveSciBench-GRN (LLM-AutoSciLab) - Edge F1: leaderboard

Metric: Edge F1 (%; directed-edge recovery against the hidden regulatory graph over the 45 ActiveSciBench-GRN tasks; LLM-AutoSciLab framework with this backbone). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o Mini72.49
2Qwen 3 32B62.7
3Qwen 3 14B51.19
4Qwen 3 4B49.56

Interactive version: theaggregate.ai/benchmark?slug=activescibench-grn-llm-autoscilab-edge-f1 · How It Works · Data refreshed daily, snapshot 2026-09-26.