Qwen 3.5 27B FP8: benchmark results
Provider: Alibaba. Released 2026-02-24. Access: Open.
Unified ELO 1732 ± 27, rank #174 of 1605 rated models, from 67 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Bulgarian Common Sense Reasoning | 83.19 | Common Sense Reasoning Average Score (%) | 100 |
| IFEval-TR - Average (loose) | 79.11 | Accuracy (%) | 100 |
| IFEval-TR - Translated (loose) | 69.94 | Accuracy (%) | 100 |
| IFEval-TR - Translated (strict) | 69.1 | Accuracy (%) | 100 |
| IFEval-TR - Turkish (loose) | 88.27 | Accuracy (%) | 100 |
| IFEval-TR - Turkish (strict) | 80 | Accuracy (%) | 100 |
| MultiGlobeQA (Gold Triples with Code Agent) | 61.6 | Exact-match accuracy (%; oracle Tier 3: the gold triples giv | 100 |
| MultiGlobeQA (KG Retrieval Agent) | 44.3 | Exact-match accuracy (%; Tier 3a: a CodeAgent writing and ex | 100 |
| MultiGlobeQA (KG and Web Agent) | 44.1 | Exact-match accuracy (%; Tier 3c: a CodeAgent with both know | 100 |
| EuroEval Catalan Common Sense Reasoning | 80.3 | Common Sense Reasoning Average Score (%) | 99.5 |
| EuroEval Croatian Common Sense Reasoning | 81.04 | Common Sense Reasoning Average Score (%) | 99.5 |
| EuroEval Danish NLU - Angry Tweets | 63.48 | Sentiment classification Score (%) | 99.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-27b-fp8 · How It Works · Data refreshed daily, snapshot 2026-09-26.