ToxReason - Liver Toxicity: leaderboard
Metric: F1 (%) for liver toxicity predicted by the model on the ToxReason test set (query chemicals with curated human CTD toxicity labels and adverse outcome pathway context, eight retrieved activation and inhibition examples from similar compounds), zero-shot, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 65 |
| 2 | GPT-4o | 60.2 |
| 3 | DeepSeek R1 Distill Llama 70B | 59.6 |
| 4 | GPT-5.1 | 58.9 |
| 5 | O3 | 58.8 |
| 6 | Qwen 2.5 14B Instruct | 58.3 |
| 7 | Llama 3.1 70B Instruct | 57.4 |
| 8 | Qwen 3 4B Instruct | 57.3 |
| 9 | Gemma 3 27B (IT) | 56.6 |
| 10 | Llama 3.1 8B Instruct | 55.7 |
| 11 | O4 Mini | 55.6 |
Interactive version: theaggregate.ai/benchmark?slug=toxreason-liver-toxicity · How It Works · Data refreshed daily, snapshot 2026-10-07.