MeSH-Rel-4K: leaderboard

Metric: Macro F1 (%; mean F1 over the broader, narrower, same-as and other relation classes for 800 held-out MeSH topic pairs; standard zero-shot prompt; 4-bit quantized open models run through KoboldAI). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1Gemma 2 9B (IT) [4bit]66.9
2Mistral 7B Instruct (v0.3) [4bit]60.3
3Phi-3.5-mini-instruct [4bit]52.9
4mistral-7B-sft-beta [4bit]41.5
5Llama 3.2 3B Instruct [4bit]25.2

Interactive version: theaggregate.ai/benchmark?slug=mesh-rel-4k · How It Works · Data refreshed daily, snapshot 2026-09-29.