tulu-2-7B — benchmark results
Provider: Allen AI. Released 2023-11-13. Access: Open.
Unified ELO 1293 ± 17, rank #1545 of 1776 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Trustworthy - Stereotype | 96.6 | Trust Score (%) | 68 |
| LLM Trustworthy - Out-of-Distribution | 69.3 | Trust Score (%) | 60 |
| LLM Trustworthy - Adversarial | 44.62 | Trust Score (%) | 44 |
| LLM Trustworthy - Adversarial Demo | 60.49 | Trust Score (%) | 44 |
| LLM Trustworthy Leaderboard | 63.56 | Average Trust Score (%) | 44 |
| Open CoT - LogiQA | 3.67 | CoT Gain (%) | 43.1 |
| Finetuning with Scientific Data Increases Hall | 62.7 | OFS (self-reported) | 41.2 |
| Open CoT - LSAT Analytical Reasoning | 3.04 | CoT Gain (%) | 41.2 |
| LLM Trustworthy - Ethics | 49 | Trust Score (%) | 40 |
| LLM Trustworthy - Fairness | 83.21 | Trust Score (%) | 36 |
| Open CoT - LSAT Reading Comprehension | 8.55 | CoT Gain (%) | 35.5 |
| Open CoT Leaderboard | 5.29 | Average CoT Gain (%) | 35.1 |
Interactive version: theaggregate.ai/model?slug=tulu-2-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.