Qwen 3 4B 2507 (Thinking): benchmark results
Qwen 3 4B 2507 evaluated with thinking enabled. Provider: Alibaba. Released 2025-07-01. Access: Open.
Unified ELO 1487 ± 1, rank #928 of 1761 rated models, from 369 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Slovak NLU - UNER SK | 73.36 | Named entity recognition Score (%) | 100 |
| EuroEval German NLU - GermEval | 78.27 | Named entity recognition Score (%) | 99.4 |
| EuroEval Dutch NLU - CoNLL NL | 80.14 | Named entity recognition Score (%) | 98.8 |
| EuroEval Italian NLU - MultiNERD IT | 87.46 | Named entity recognition Score (%) | 98.8 |
| EuroEval Spanish NLU | 60.1 | NLU Average Score (%) | 97.3 |
| EuroEval Latvian NLU - Fullstack NER LV | 73.2 | Named entity recognition Score (%) | 97 |
| EuroEval Ukrainian NLU - NER UK | 79.05 | Named entity recognition Score (%) | 97 |
| MERA Code - ruHumanEval | 78.11 | pass@1 (%) | 96.9 |
| EuroEval Norwegian NLU - NorNE NB | 86.13 | Named entity recognition Score (%) | 96.7 |
| MERA - ruHumanEval | 77.44 | pass@1 (%) | 96.6 |
| EuroEval Danish NLU - Dansk | 68.78 | Named entity recognition Score (%) | 96.3 |
| EuroEval French NLU - Eltec | 73.38 | Named entity recognition Score (%) | 96 |
Interactive version: theaggregate.ai/model?slug=qwen-3-4b-2507-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.