Sommerfugl-31B-v2
Sommerfugl ("butterfly") is a family of Norwegian language models by oolabs.no.
Built with Gemma. This model is a finetune of google/gemma-4-31B-it, modified by oolabs.no. Use is governed by the Gemma Terms of Use.
Why
Google's instruction tuning of Gemma 4 damages its Norwegian: on the National Library of Norway's leaderboard, gemma-4-31B-it ranks in the top 3 on eleven categories but collapses to ~rank 48 on Norwegian language knowledge and ~rank 51 on Norwegian summarization. Sommerfugl-31B-v2 repairs both while retaining the base's strengths.
What changed from v1
Sommerfugl-31B-v1 often answered simple closed-book questions ("Who wrote Et dukkehjem?") with "Det er ikke oppgitt…" ("that is not stated") instead of the answer. v2 answers them directly, and improves TruthfulQA, NCB and NorQuAD over v1.
Results
Evaluated under the National Library's published leaderboard protocol (lm-evaluation-harness + NorEval; cloze MC, chat template, best over prompt versions × {0,5}-shot, greedy decoding). Gemma column: NB's own leaderboard row. Sommerfugl and Borealis 2 (NbAiLab's newest flagship) evaluated by us under the identical protocol.
| task (metric) | gemma-4-31B-it | Sommerfugl-31B-v2 | Borealis 2 preview |
|---|---|---|---|
| Language knowledge: NoCoLA (acc) | 0.828 | 0.866 | 0.834 |
| Language knowledge: NCB (acc) | 0.755 | 0.849 | 0.774 |
| Summaries: NorSumm nob / nno (bleu) | 4.08 / 3.37 | 6.82 / 7.02 | 6.62 / 5.15 |
| Summarize on request (bleu) | 0.98 | 4.54 | 1.75 |
| Rewrite (bleu) | 0.03 | 3.96 | 0.11 |
| Idioms nob / nno (fscore) | 0.085 / 0.141 | 0.441 / 0.507 | 0.121 / 0.160 |
| Grammar correction (exact match) | 0.294 | 0.405 | 0.342 |
| Reading comp.: NorQuAD (f1) | 0.445 | 0.881 | 0.799 |
| Belebele (acc) | 0.938 | 0.940 | 0.904 |
| OpenbookQA / CommonsenseQA nob (acc) | 0.968 / 0.860 | 0.960 / 0.838 | 0.955 / 0.808 |
| NoReC sentiment (acc) | 0.910 | 0.935 | 0.918 |
| TruthfulQA mc nob / nno (acc) | 0.850 / 0.930 | 0.738 / 0.789 | 0.811 / 0.842 |
| TruthfulQA gen nob / nno (rougeL acc) | 0.566 / 0.608 | 0.558 / 0.552 | 0.561 / 0.592 |
| Translation en→nb / en→nn (bleu) | 59.5 / 48.5 | 61.3 / 49.3 | 59.6 / 48.4 |
| Translation nb→en / nn→en (bleu) | 60.4 / 60.4 | 62.4 / 59.5 | 63.0 / 61.3 |
| MMLU English (acc) | 0.831 | 0.834 | 0.550 |
| MMLU Norwegian (acc) | 0.859 | 0.797¹ | — |
¹ 0-shot only; the 5-shot run was not completed.
Head-to-head vs Borealis 2 across all 25 protocol-complete values (paired rows counted separately, plus the Nynorsk OpenbookQA/CommonsenseQA variants): 19–6. Borealis 2 leads on TruthfulQA and on translation into English.
Known limitations (reported deliberately)
- TruthfulQA (mc) is below the base -it (−0.11 nob / −0.14 nno): supervised finetuning erodes part of the base's truthfulness calibration. Verify facts, figures and sources in its answers before relying on them.
- MMLU-nb is below the base -it (0.797 0-shot vs 0.859).
- Translation into English trails Borealis 2 (−0.6 / −1.8 bleu).
- Not safety-evaluated: NB's safety sets are among the tasks we could not run. Put your own filters and policies in front of it before exposing it to end users.
- Not evaluated: NB's unpublished internal tasks (translated ARC/GSM8K/IFEval/GoldenSwag, safety sets — ~8% of their suite); exclusions applied identically to every model we compare.
- We publish no overall/aggregate score: that number belongs to the National Library's own evaluation pipeline.
Evaluation integrity
Training data was quality-gated and decontaminated against all 19 Norwegian evaluation datasets used above (all splits), so the reported numbers are not inflated by train/test overlap.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("oolabs/sommerfugl-31b-v2")
model = AutoModelForCausalLM.from_pretrained("oolabs/sommerfugl-31b-v2", dtype="bfloat16", device_map="auto")
msgs = [{"role": "user", "content": "Skriv et kort sammendrag av teksten under."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
- Downloads last month
- 29