MeSH-Rel-4K (CoT Two-Way): leaderboard
Metric: Macro F1 (%; mean F1 over the broader, narrower, same-as and other relation classes for 800 held-out MeSH topic pairs; two-stage chain-of-thought prompt run in both topic orders and reconciled by fixed rules; 4-bit quantized open models run through KoboldAI). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemma 2 9B (IT) [4bit] | 71.6 |
| 2 | Mistral 7B Instruct (v0.3) [4bit] | 70.3 |
| 3 | Phi-3.5-mini-instruct [4bit] | 57.6 |
| 4 | mistral-7B-sft-beta [4bit] | 48.6 |
| 5 | Llama 3.2 3B Instruct [4bit] | 25.7 |
Interactive version: theaggregate.ai/benchmark?slug=mesh-rel-4k-cot-two-way · How It Works · Data refreshed daily, snapshot 2026-09-29.