gpt-sw3-126m-instruct: benchmark results

Provider: AI Sweden. Released 2023-04-28. Access: Open.

Unified ELO 1191 ± 20, rank #2912 of 2928 rated models, from 97 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Icelandic NLU - ScaLA IS2.31Linguistic acceptability Score (%)45.8
Open LLM Leaderboard v1 - TruthfulQA MC242.65MC2 (%) (0-shot)21.2
EuroEval Icelandic Common Sense Reasoning0.59Common Sense Reasoning Average Score (%)19.8
EuroEval Icelandic NLU - MIM-GOLD NER17.85Named entity recognition Score (%)18.7
EuroEval Danish NLU - Angry Tweets17.37Sentiment classification Score (%)16.3
EuroEval Faroese NLU - ScaLA FO0Linguistic acceptability Score (%)15.9
EuroEval Norwegian8.13Average Score (%)15.8
Open LLM Leaderboard v1 - GSM8K0.99Accuracy (%) (5-shot)15.1
EuroEval Dutch Common Sense Reasoning0.69Common Sense Reasoning Average Score (%)14.8
EuroEval Faroese NLU - FONE29.27Named entity recognition Score (%)14.7
EuroEval Norwegian NLU - Norquad9.58Reading comprehension Score (%)14.4
EuroEval Icelandic NLU - Hotter and Colder Sentiment3.39Sentiment classification Score (%)14

Interactive version: theaggregate.ai/model?slug=gpt-sw3-126m-instruct · How It Works · Data refreshed daily, snapshot 2026-09-23.