Sommerfugl-31B-v2

Sommerfugl ("butterfly") is a family of Norwegian language models by oolabs.no.

Built with Gemma. This model is a finetune of google/gemma-4-31B-it, modified by oolabs.no. Use is governed by the Gemma Terms of Use.

Why

Google's instruction tuning of Gemma 4 damages its Norwegian: on the National Library of Norway's leaderboard, gemma-4-31B-it ranks in the top 3 on eleven categories but collapses to ~rank 48 on Norwegian language knowledge and ~rank 51 on Norwegian summarization. Sommerfugl-31B-v2 repairs both while retaining the base's strengths.

What changed from v1

Sommerfugl-31B-v1 often answered simple closed-book questions ("Who wrote Et dukkehjem?") with "Det er ikke oppgitt…" ("that is not stated") instead of the answer. v2 answers them directly, and improves TruthfulQA, NCB and NorQuAD over v1.

Results

Evaluated under the National Library's published leaderboard protocol (lm-evaluation-harness + NorEval; cloze MC, chat template, best over prompt versions × {0,5}-shot, greedy decoding). Gemma column: NB's own leaderboard row. Sommerfugl and Borealis 2 (NbAiLab's newest flagship) evaluated by us under the identical protocol.

task (metric) gemma-4-31B-it Sommerfugl-31B-v2 Borealis 2 preview
Language knowledge: NoCoLA (acc) 0.828 0.866 0.834
Language knowledge: NCB (acc) 0.755 0.849 0.774
Summaries: NorSumm nob / nno (bleu) 4.08 / 3.37 6.82 / 7.02 6.62 / 5.15
Summarize on request (bleu) 0.98 4.54 1.75
Rewrite (bleu) 0.03 3.96 0.11
Idioms nob / nno (fscore) 0.085 / 0.141 0.441 / 0.507 0.121 / 0.160
Grammar correction (exact match) 0.294 0.405 0.342
Reading comp.: NorQuAD (f1) 0.445 0.881 0.799
Belebele (acc) 0.938 0.940 0.904
OpenbookQA / CommonsenseQA nob (acc) 0.968 / 0.860 0.960 / 0.838 0.955 / 0.808
NoReC sentiment (acc) 0.910 0.935 0.918
TruthfulQA mc nob / nno (acc) 0.850 / 0.930 0.738 / 0.789 0.811 / 0.842
TruthfulQA gen nob / nno (rougeL acc) 0.566 / 0.608 0.558 / 0.552 0.561 / 0.592
Translation en→nb / en→nn (bleu) 59.5 / 48.5 61.3 / 49.3 59.6 / 48.4
Translation nb→en / nn→en (bleu) 60.4 / 60.4 62.4 / 59.5 63.0 / 61.3
MMLU English (acc) 0.831 0.834 0.550
MMLU Norwegian (acc) 0.859 0.797¹ —

¹ 0-shot only; the 5-shot run was not completed.

Head-to-head vs Borealis 2 across all 25 protocol-complete values (paired rows counted separately, plus the Nynorsk OpenbookQA/CommonsenseQA variants): 19–6. Borealis 2 leads on TruthfulQA and on translation into English.

Known limitations (reported deliberately)

  • TruthfulQA (mc) is below the base -it (−0.11 nob / −0.14 nno): supervised finetuning erodes part of the base's truthfulness calibration. Verify facts, figures and sources in its answers before relying on them.
  • MMLU-nb is below the base -it (0.797 0-shot vs 0.859).
  • Translation into English trails Borealis 2 (−0.6 / −1.8 bleu).
  • Not safety-evaluated: NB's safety sets are among the tasks we could not run. Put your own filters and policies in front of it before exposing it to end users.
  • Not evaluated: NB's unpublished internal tasks (translated ARC/GSM8K/IFEval/GoldenSwag, safety sets — ~8% of their suite); exclusions applied identically to every model we compare.
  • We publish no overall/aggregate score: that number belongs to the National Library's own evaluation pipeline.

Evaluation integrity

Training data was quality-gated and decontaminated against all 19 Norwegian evaluation datasets used above (all splits), so the reported numbers are not inflated by train/test overlap.

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("oolabs/sommerfugl-31b-v2")
model = AutoModelForCausalLM.from_pretrained("oolabs/sommerfugl-31b-v2", dtype="bfloat16", device_map="auto")
msgs = [{"role": "user", "content": "Skriv et kort sammendrag av teksten under."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Downloads last month
29
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oolabs/sommerfugl-31b-v2

Finetuned
(293)
this model
Quantizations
3 models