MeSH-Rel-4K (CoT Two-Way): leaderboard

Metric: Macro F1 (%; mean F1 over the broader, narrower, same-as and other relation classes for 800 held-out MeSH topic pairs; two-stage chain-of-thought prompt run in both topic orders and reconciled by fixed rules; 4-bit quantized open models run through KoboldAI). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.

Top models

#ModelScore
1Gemma 2 9B (IT) [4bit]71.6
2Mistral 7B Instruct (v0.3) [4bit]70.3
3Phi-3.5-mini-instruct [4bit]57.6
4mistral-7B-sft-beta [4bit]48.6
5Llama 3.2 3B Instruct [4bit]25.7

Interactive version: theaggregate.ai/benchmark?slug=mesh-rel-4k-cot-two-way · How It Works · Data refreshed daily, snapshot 2026-09-29.