ActiveSciBench-Chem (LLM-AutoSciLab): leaderboard

Metric: Symbolic accuracy (%; share of the 57 enzyme-kinetics tasks whose recovered rate law an LLM judge rates equivalent to the hidden law up to fitted constants; budget-limited active experiments driven by the LLM-AutoSciLab framework with this backbone). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4o Mini35.09
2Qwen 3 32B25.56
3Qwen 3 14B23.88
4Qwen 3 4B12.92

Interactive version: theaggregate.ai/benchmark?slug=activescibench-chem-llm-autoscilab · How It Works · Data refreshed daily, snapshot 2026-09-26.