Qwen3 235B A22B FP8 Throughput — benchmark results
Qwen 3 235B A22B evaluated in FP8 quantization on a throughput-optimized serving profile. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1583 ± 39, rank #523 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety XSTest | 98.8 | LM Evaluated Safety score (%) | 98.8 |
| HELM Capabilities - Omni-MATH | 54.8 | Acc | 88 |
| HELM Capabilities - MMLU-Pro | 81.7 | COT correct | 82 |
| HELM Safety BBQ | 96.6 | BBQ accuracy (%) | 79.7 |
| HELM Capabilities - WildBench | 82.78 | WB Score | 74 |
| HELM Capabilities - GPQA | 62.33 | COT correct | 70 |
| HELM Safety Anthropic Red Team | 99 | LM Evaluated Safety score (%) | 53.5 |
| HELM Safety SimpleSafetyTests | 98.5 | LM Evaluated Safety score (%) | 46.5 |
| HELM Capabilities - IFEval | 81.58 | IFEval Strict Acc | 42 |
| HELM Safety | 90 | Mean score (self-reported) | 41.3 |
| HELM AIR-Bench | 56 | Refusal Rate (%) | 27.9 |
| HELM Safety HarmBench | 56.9 | LM Evaluated Safety score (%) | 20.9 |
Interactive version: theaggregate.ai/model?slug=qwen3-235b-a22b-fp8-throughput · How the rankings work · Data refreshed daily, snapshot 2026-07-22.