dr-tulu: benchmark results

Provider: Other.

Unified ELO 1703 ± 22, rank #188 of 1636 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Mr Dre - Content Feedback Break Rate23.5Break rate (%; share of checklist criteria covered in the tu100
Mr Dre - Format Feedback Break Rate14Break rate (%; share of checklist criteria covered in the tu100
Mr Dre - Format Feedback Incorporation Rate94.9Incorporation rate (%; turn-2 revision after format feedback75
SAGE (Literature Search) - Short-Form (Computer Science)36Exact match (%, 0-100): the target paper appears in the answ60
SAGE (Literature Search) - Short-Form (Healthcare)58Exact match (%, 0-100): the target paper appears in the answ60
SAGE (Literature Search) - Short-Form (Natural Science)44.7Exact match (%, 0-100): the target paper appears in the answ60
Mr Dre - Self-Reflection Break Rate11.4Break rate (%; share of checklist criteria covered in the tu50
SAGE (Literature Search) - Open-Ended (Healthcare)17.2Weighted recall (%, 0-100) of the ground-truth papers found 40
SAGE (Literature Search) - Open-Ended (Humanities)14.1Weighted recall (%, 0-100) of the ground-truth papers found 40
SAGE (Literature Search) - Open-Ended (Natural Science)20.2Weighted recall (%, 0-100) of the ground-truth papers found 40
SAGE (Literature Search) - Short-Form (Humanities)49.3Exact match (%, 0-100): the target paper appears in the answ40
DeepResearch Bench - Instruction Following49.56Score (%)36.4

Interactive version: theaggregate.ai/model?slug=dr-tulu · How It Works · Data refreshed daily, snapshot 2026-10-11.