MeSH-Rel-4K: leaderboard
Metric: Macro F1 (%; mean F1 over the broader, narrower, same-as and other relation classes for 800 held-out MeSH topic pairs; standard zero-shot prompt; 4-bit quantized open models run through KoboldAI). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 2 9B (IT) [4bit] | 66.9 |
| 2 | Mistral 7B Instruct (v0.3) [4bit] | 60.3 |
| 3 | Phi-3.5-mini-instruct [4bit] | 52.9 |
| 4 | mistral-7B-sft-beta [4bit] | 41.5 |
| 5 | Llama 3.2 3B Instruct [4bit] | 25.2 |
Interactive version: theaggregate.ai/benchmark?slug=mesh-rel-4k · How It Works · Data refreshed daily, snapshot 2026-09-29.