Seed 2.0 Pro (Thinking): benchmark results
Provider: ByteDance. Released 2026-02-14. Access: API.
Unified ELO 1661 ± 1, rank #280 of 3078 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLMEval-Logic Base | 75.5 | Accuracy (%) | 100 |
| VitaBench 2.0 | 47.4 | Avg@4 Full Context (self-reported) | 92 |
| LLMEval-Logic Formalization Fixed | 56.9 | Accuracy (%) | 65.4 |
| LLMEval-Logic Formalization Free | 35.8 | Accuracy (%) | 53.8 |
| ComboShoppingBench - Response Quality | 81.8 | Pass rate (%; LLM-judged) | 47.6 |
| LLMEval-Logic Hard Sub-Q | 63.3 | Accuracy (%) | 38.5 |
| ComboShoppingBench - Overall Success | 17.5 | Pass rate (%; all judged and rule-based checks) | 38.1 |
| ComboShoppingBench - Budget Compliance | 72.9 | Pass rate (%) | 33.3 |
| ComboShoppingBench - Coupon-ID Validity | 96.2 | Pass rate (%) | 33.3 |
| LLMEval-Logic Hard | 20.4 | Accuracy (%) | 30.8 |
| ComboShoppingBench - Claim Faithfulness | 69.4 | Pass rate (%; LLM-judged) | 28.6 |
| ComboShoppingBench - Coupon Optimality | 56 | Pass rate (%) | 23.8 |
Interactive version: theaggregate.ai/model?slug=seed-2-0-pro-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.