OpenTumorBoard - Board Simulation: leaderboard
Metric: Conclusion alignment (1-5) of a simulated multidisciplinary discussion's consensus with the real tumor board's conclusion, mean over 184 test cases (malformed outputs score 1; judged by Qwen3.8-27B); higher is better. Source: arxiv.org. Saturation forecast: Around 2035. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.7 Flash | 2.78 |
| 2 | Grok 4.6 | 2.58 |
| 3 | Qwen 3.8 Max | 2.57 |
| 4 | GPT-5.6 Sol | 2.53 |
| 5 | Claude Opus 5 (Thinking) | 2.53 |
| 6 | Ministral 3 14B | 2.2 |
| 7 | MedGemma-27B-IT | 2.15 |
| 8 | Llama 4 Scout | 2.1 |
Interactive version: theaggregate.ai/benchmark?slug=opentumorboard-board-simulation · How It Works · Data refreshed daily, snapshot 2026-09-29.