dr-tulu: benchmark results
Provider: Other.
Unified ELO 1703 ± 22, rank #188 of 1636 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Mr Dre - Content Feedback Break Rate | 23.5 | Break rate (%; share of checklist criteria covered in the tu | 100 |
| Mr Dre - Format Feedback Break Rate | 14 | Break rate (%; share of checklist criteria covered in the tu | 100 |
| Mr Dre - Format Feedback Incorporation Rate | 94.9 | Incorporation rate (%; turn-2 revision after format feedback | 75 |
| SAGE (Literature Search) - Short-Form (Computer Science) | 36 | Exact match (%, 0-100): the target paper appears in the answ | 60 |
| SAGE (Literature Search) - Short-Form (Healthcare) | 58 | Exact match (%, 0-100): the target paper appears in the answ | 60 |
| SAGE (Literature Search) - Short-Form (Natural Science) | 44.7 | Exact match (%, 0-100): the target paper appears in the answ | 60 |
| Mr Dre - Self-Reflection Break Rate | 11.4 | Break rate (%; share of checklist criteria covered in the tu | 50 |
| SAGE (Literature Search) - Open-Ended (Healthcare) | 17.2 | Weighted recall (%, 0-100) of the ground-truth papers found | 40 |
| SAGE (Literature Search) - Open-Ended (Humanities) | 14.1 | Weighted recall (%, 0-100) of the ground-truth papers found | 40 |
| SAGE (Literature Search) - Open-Ended (Natural Science) | 20.2 | Weighted recall (%, 0-100) of the ground-truth papers found | 40 |
| SAGE (Literature Search) - Short-Form (Humanities) | 49.3 | Exact match (%, 0-100): the target paper appears in the answ | 40 |
| DeepResearch Bench - Instruction Following | 49.56 | Score (%) | 36.4 |
Interactive version: theaggregate.ai/model?slug=dr-tulu · How It Works · Data refreshed daily, snapshot 2026-10-11.