gpt-sw3-126m-instruct: benchmark results
Provider: AI Sweden. Released 2023-04-28. Access: Open.
Unified ELO 1191 ± 20, rank #2912 of 2928 rated models, from 97 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Icelandic NLU - ScaLA IS | 2.31 | Linguistic acceptability Score (%) | 45.8 |
| Open LLM Leaderboard v1 - TruthfulQA MC2 | 42.65 | MC2 (%) (0-shot) | 21.2 |
| EuroEval Icelandic Common Sense Reasoning | 0.59 | Common Sense Reasoning Average Score (%) | 19.8 |
| EuroEval Icelandic NLU - MIM-GOLD NER | 17.85 | Named entity recognition Score (%) | 18.7 |
| EuroEval Danish NLU - Angry Tweets | 17.37 | Sentiment classification Score (%) | 16.3 |
| EuroEval Faroese NLU - ScaLA FO | 0 | Linguistic acceptability Score (%) | 15.9 |
| EuroEval Norwegian | 8.13 | Average Score (%) | 15.8 |
| Open LLM Leaderboard v1 - GSM8K | 0.99 | Accuracy (%) (5-shot) | 15.1 |
| EuroEval Dutch Common Sense Reasoning | 0.69 | Common Sense Reasoning Average Score (%) | 14.8 |
| EuroEval Faroese NLU - FONE | 29.27 | Named entity recognition Score (%) | 14.7 |
| EuroEval Norwegian NLU - Norquad | 9.58 | Reading comprehension Score (%) | 14.4 |
| EuroEval Icelandic NLU - Hotter and Colder Sentiment | 3.39 | Sentiment classification Score (%) | 14 |
Interactive version: theaggregate.ai/model?slug=gpt-sw3-126m-instruct · How It Works · Data refreshed daily, snapshot 2026-09-23.