ToxReason - Mechanistic Reasoning: leaderboard
Metric: Overall quality (0-10) of the step-wise mechanistic explanation linking molecular initiating events to the organ-level adverse outcome, scored 0 to 10 by a Claude Sonnet 4.5 judge against the reference adverse outcome pathway, on the ToxReason test set (query chemicals with curated human CTD toxicity labels and adverse outcome pathway context, eight retrieved activation and inhibition examples from similar compounds), zero-shot, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.1 | 5.52 |
| 2 | GPT-5 | 5.42 |
| 3 | O3 | 5.33 |
| 4 | GPT-4o | 4.96 |
| 5 | Gemma 3 27B (IT) | 4.96 |
| 6 | O4 Mini | 4.95 |
| 7 | Llama 3.1 70B Instruct | 4.65 |
| 8 | Qwen 2.5 14B Instruct | 4.53 |
| 9 | Qwen 3 4B Instruct | 4.52 |
| 10 | DeepSeek R1 Distill Llama 70B | 4.49 |
| 11 | Llama 3.1 8B Instruct | 3.28 |
Interactive version: theaggregate.ai/benchmark?slug=toxreason-mechanistic-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.