ActiveSciBench-GRN (LLM-AutoSciLab): leaderboard

Metric: Exact graph accuracy (%; share of the 45 gene-regulatory-network tasks whose hidden signed causal graph is recovered exactly; budget-limited perturbation experiments driven by the LLM-AutoSciLab framework with this backbone). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o Mini31.11
2Qwen 3 32B25.88
3Qwen 3 14B22.34
4Qwen 3 4B21.78

Interactive version: theaggregate.ai/benchmark?slug=activescibench-grn-llm-autoscilab · How It Works · Data refreshed daily, snapshot 2026-09-26.