TreeProbe (One-Shot): leaderboard

Metric: MCQ accuracy (%; unweighted macro mean over the ten subtasks; over TreeProbe's 3,289 expert-adjudicated Tibetan-medicine multiple-choice items in ten subtasks across three diagnostic roots, option order randomized, English prompt with Tibetan-script content; one-shot prompting with a fixed per-subtask exemplar). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Gemini 3.1 Flash Lite65.9
2GPT-5.465.47
3Qwen 3.5 397B A17B64.27
4GLM-4.757.47
5Claude Sonnet 4.645.22

Interactive version: theaggregate.ai/benchmark?slug=treeprobe-one-shot · How It Works · Data refreshed daily, snapshot 2026-09-26.