gpt-sw3-6.7B-v2-instruct: benchmark results
Provider: AI Sweden. Released 2023-04-28. Access: Open.
Unified ELO 1280 ± 21, rank #2754 of 2928 rated models, from 36 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Norwegian NLU - Norec | 34.39 | Sentiment classification Score (%) | 35.6 |
| EuroEval Danish NLU - ScaLA DA | 10.99 | Linguistic acceptability Score (%) | 33.9 |
| EuroEval Norwegian Knowledge - Idioms NO | 0.75 | MCC (x100) | 33.6 |
| EuroEval Swedish NLU - ScaLA SV | 10.92 | Linguistic acceptability Score (%) | 32 |
| EuroEval Norwegian NLU - ScaLA NN | 5.11 | Linguistic acceptability Score (%) | 31.6 |
| EuroEval Icelandic Knowledge | 3.28 | Knowledge Average Score (%) | 31.3 |
| EuroEval Danish Common Sense Reasoning | 11.08 | Common Sense Reasoning Average Score (%) | 30.1 |
| EuroEval Swedish NLU - Multi Wiki QA SV | 38.22 | Reading comprehension Score (%) | 29.2 |
| EuroEval Swedish Common Sense Reasoning | 10.9 | Common Sense Reasoning Average Score (%) | 29 |
| EuroEval Norwegian Knowledge | 4.46 | Knowledge Average Score (%) | 26.6 |
| Open LLM Leaderboard v1 - GSM8K | 6.37 | Accuracy (%) (5-shot) | 25.5 |
| EuroEval Norwegian Knowledge - NRK Quiz QA | 8.16 | MCC (x100) | 25.2 |
Interactive version: theaggregate.ai/model?slug=gpt-sw3-6-7b-v2-instruct · How It Works · Data refreshed daily, snapshot 2026-09-23.