TreeProbe (Zero-Shot): leaderboard

Metric: MCQ accuracy (%; unweighted macro mean over the ten subtasks; over TreeProbe's 3,289 expert-adjudicated Tibetan-medicine multiple-choice items in ten subtasks across three diagnostic roots, option order randomized, English prompt with Tibetan-script content; zero-shot prompting). Source: arxiv.org. Saturation forecast: Around June 2027. 6 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B60.55
2Gemini 3.1 Flash Lite59.89
3GPT-5.458.1
4GLM-4.752.11
5Claude Sonnet 4.640.61

Interactive version: theaggregate.ai/benchmark?slug=treeprobe-zero-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.