tulu-2-13B — benchmark results
Provider: Allen AI. Released 2023-11-13. Access: Open.
Unified ELO 1283 ± 15, rank #1572 of 1776 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Trustworthy - Fairness | 97.9 | Trust Score (%) | 76 |
| LLM Trustworthy - Adversarial Demo | 71.17 | Trust Score (%) | 72 |
| LLM Trustworthy - Out-of-Distribution | 70.17 | Trust Score (%) | 64 |
| LLM Trustworthy - Stereotype | 89.33 | Trust Score (%) | 56 |
| LLM Trustworthy Leaderboard | 66.51 | Average Trust Score (%) | 52 |
| LLM Trustworthy - Toxicity | 44.8 | Trust Score (%) | 48 |
| LLM Trustworthy - Privacy | 78.9 | Trust Score (%) | 44 |
| Open CoT - LogiQA 2 | 5.47 | CoT Gain (%) | 42 |
| BiGGen-Bench | 3.21 | Average Score (1-5) | 41.2 |
| Open CoT - LSAT Logical Reasoning | 8.63 | CoT Gain (%) | 39.7 |
| Open CoT - LSAT Reading Comprehension | 7.81 | CoT Gain (%) | 32.8 |
| Open CoT Leaderboard | 5.03 | Average CoT Gain (%) | 32.4 |
Interactive version: theaggregate.ai/model?slug=tulu-2-13b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.