ToxReason - Biological Fidelity: leaderboard
Metric: Biological fidelity (0-10) of the mechanistic explanation, scored 0 to 10 by a Claude Sonnet 4.5 judge against the reference adverse outcome pathway, on the ToxReason test set (query chemicals with curated human CTD toxicity labels and adverse outcome pathway context, eight retrieved activation and inhibition examples from similar compounds), zero-shot, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.1 | 6.38 |
| 2 | GPT-5 | 6.18 |
| 3 | O3 | 5.92 |
| 4 | O4 Mini | 5.49 |
| 5 | Gemma 3 27B (IT) | 5.45 |
| 6 | GPT-4o | 5.31 |
| 7 | Qwen 3 4B Instruct | 5.1 |
| 8 | Llama 3.1 70B Instruct | 5.02 |
| 9 | Qwen 2.5 14B Instruct | 5 |
| 10 | DeepSeek R1 Distill Llama 70B | 4.88 |
| 11 | Llama 3.1 8B Instruct | 3.74 |
Interactive version: theaggregate.ai/benchmark?slug=toxreason-biological-fidelity · How It Works · Data refreshed daily, snapshot 2026-10-07.