You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Sunflower-Qwen3.8-27B

Sunbird AI's open model for 68 African languages, with the deepest coverage of the 31 languages of Uganda. It is the strongest open-weight model available for translation between English and African languages, it matches Gemini 3.1 Pro when translating into African languages while beating GPT-5.1 and Claude Sonnet 4.5 by a wide margin, and it reasons adaptively: with thinking enabled it decides per prompt whether a reasoning trace is worth it, so maths and knowledge questions get one and translation stays fast.

Model Base Use it for
Sunflower-Qwen3.8-27B (this model) Qwen3.8-27B Best quality across translation, comprehension and reasoning
Sunflower-Qwen3.5-9B Qwen3.5-9B Translation on smaller GPUs, within 0.4 chrF of the 27B

Highlights

  • Translation, 68 African languages: chrF 46.2, against 32.3 for the best other open model, 38.9 for Claude Sonnet 4.5 and 38.0 for GPT-5.1.
  • Into African languages, the hard direction, Sunflower is level with Gemini 3.1 Pro across the 68-language set (41.6 vs 41.8) and ahead of it by 4.4 chrF on the 31 Ugandan languages (39.2 vs 34.8), producing roughly one thirtieth of the output tokens per sentence.
  • FLORES-200: best open model on the public benchmark (chrF 53.5 vs 49.1), ahead of GPT-5.1.
  • SAHARA: 42.15, first among open-weight models, sixth overall behind five closed models.
  • Comprehension and reasoning: best open model on four of six AfroBench-Lite tasks. With thinking on, AfriMGSM rises from 40.3 to 60.5 while translation quality is unchanged.
  • The 9B stands on its own: chrF 45.8 on the 68 languages, ahead of every other open model, including ones three times its size.

Evaluation

Scores are chrF or accuracy on a 0 to 100 scale, thinking off unless stated. Every number comes from the public evaluation code at github.com/SunbirdAI/sunflower, with the same sentences, prompts and output budgets for every model. Sunflower's own starting checkpoints are shown alongside, so the gain from training can be read directly.

Translation

Mean chrF over both directions.

Model 68 African languages¹ 31 Ugandan languages² FLORES-200, 14 languages³
Gemini 3.1 Pro⁴ 48.5 42.1 58.1
Sunflower-Qwen3.8-27B 46.2 42.4 53.5
Sunflower-Qwen3.5-9B 45.8 42.0 52.0
Claude Sonnet 4.5 38.9 30.6 n/a⁵
GPT-5.1 38.0 30.8 51.4
AfriqueQwen3.5-9B Instruct v1 32.3 23.3 49.1
AfriqueQwen3.5-4B Instruct v1 31.1 22.3 47.9
Qwen3.8-27B (starting point of the 27B) 29.7 23.0 44.9
Qwen3.5-27B 29.4 21.9 42.9
Qwen3.5-9B (starting point of the 9B) 25.7 20.1 38.0

By direction. Sunflower's lead is largest when translating into African languages, the harder direction and the one most users need.

Model 68 languages, eng→xx 68 languages, xx→eng Ugandan, eng→xx Ugandan, xx→eng
Gemini 3.1 Pro⁴ 41.8 55.2 34.8 49.4
Sunflower-Qwen3.8-27B 41.6 50.8 39.2 45.7
Sunflower-Qwen3.5-9B 42.1 49.6 39.5 44.5
Claude Sonnet 4.5 35.3 42.5 27.1 34.1
GPT-5.1 33.3 42.7 26.5 35.0
AfriqueQwen3.5-9B Instruct v1 28.8 35.8 22.4 24.3
AfriqueQwen3.5-4B Instruct v1 28.2 34.1 21.4 23.2

¹ Sunbird/salt-69 test set: 68 African languages plus French, 100 sentences per language. ² Sunflower translation eval, test split: 99 sentences per language. ³ FLORES-200 devtest for the 14 AfroBench-Lite languages, 5-shot, 100 sentences per direction. ⁴ Gemini 3.1 Pro runs with reasoning on, which cannot be disabled, at roughly 30 times Sunflower's output tokens per sentence. ⁵ Withheld: a contamination probe found Claude Sonnet 4.5 reproducing the public FLORES English references verbatim. Its eng→xx score, 49.0, is valid and just below Sunflower's 49.3.

Closed models were run through their APIs with a neutral translator system prompt; Sunflower used its own. All models received the same instruction, greedy decoding and a 100-token output budget.

Translation by language

Quality tracks how much training data exists for each language. Tiers use Sunflower-27B's mean chrF on the 68-language set.

  • Good (50 and above), 30 languages: Afrikaans, Zulu, Chichewa, Swahili, Xhosa, Somali, Nigerian Pidgin, Lingala, Sotho, Malagasy, Igbo, Amharic, Hausa, Yoruba, Kinyarwanda, Tswana, Luganda, Shona, Ndebele, Kirundi, Runyoro, Ewe, Luo, Lusoga, Acholi, Oromo, Runyankole, Kikuyu, Rukiga, Akan.
  • Moderate (35 to 50), 19 languages: Rutooro, Bemba, Ruruuli, Bambara, Lango, Lugungu, Lugwere, Ateso, Kumam, Lugbara, Wolof, Lumasaba, Kabyle, Lunyole, Jopadhola, Bari, Lubwisi, Samia, Rukonjo.
  • Basic (below 35), 19 languages: Alur, Karamojong, Dagbani, Kakwa, Berber, Ma'di, Fulani, Kwamba, Aringa, Luhya, Lendu, Dinka, Dagaare, Kupsabiny, Pokot, Ik, Kalenjin, Kanuri, Ikposo. Have a speaker check output in these languages before relying on it.

French is supported as a high-resource reference language.

Scores for all 69 languages (chrF, Sunflower-Qwen3.8-27B)
Language Code eng→xx xx→eng Mean
Afrikaans afr 82.9 84.6 83.8
Zulu zul 68.4 79.7 74.0
Chichewa nya 69.4 73.9 71.7
French fra 68.7 73.0 70.9
Swahili swa 68.1 73.2 70.7
Xhosa xho 65.1 75.5 70.3
Somali som 59.7 76.8 68.2
Nigerian Pidgin pcm 57.1 78.4 67.8
Lingala lin 60.8 71.7 66.2
Sotho sot 61.4 70.9 66.1
Malagasy mlg 61.1 70.6 65.9
Igbo ibo 61.1 70.2 65.6
Amharic amh 50.3 77.5 63.9
Hausa hau 57.7 66.1 61.9
Yoruba yor 51.5 71.0 61.3
Kinyarwanda kin 59.6 62.7 61.2
Tswana tsn 57.2 64.8 61.0
Luganda lug 58.5 63.3 60.9
Shona sna 52.5 63.5 58.0
Ndebele nbl 47.4 68.5 57.9
Kirundi run 50.6 64.7 57.6
Runyoro nyo 52.0 61.4 56.7
Ewe ewe 52.0 60.2 56.1
Luo luo 52.2 57.2 54.7
Lusoga xog 46.0 61.7 53.9
Acholi ach 51.2 55.0 53.1
Oromo orm 46.5 58.4 52.4
Runyankole nyn 49.5 55.4 52.4
Kikuyu kik 40.2 62.6 51.4
Rukiga cgg 47.2 54.9 51.1
Akan aka 41.9 58.6 50.2
Rutooro ttj 43.8 55.9 49.8
Bemba bem 46.3 52.0 49.2
Ruruuli ruc 40.4 57.8 49.1
Bambara bam 39.3 53.9 46.6
Lango laj 38.8 51.6 45.2
Lugungu rub 39.2 50.6 44.9
Lugwere gwr 39.0 50.8 44.9
Ateso teo 41.1 48.1 44.6
Kumam kdi 39.3 47.7 43.5
Lugbara lgg 41.2 45.7 43.4
Wolof wol 34.4 49.1 41.7
Lumasaba myx 38.4 44.5 41.4
Kabyle kab 33.4 49.1 41.2
Lunyole nuj 36.6 44.3 40.5
Jopadhola adh 34.9 41.4 38.2
Bari bfa 34.4 40.8 37.6
Lubwisi tlj 34.1 41.0 37.5
Samia lsm 35.8 38.8 37.3
Rukonjo koo 36.2 37.4 36.8
Alur alz 29.2 39.0 34.1
Karamojong kdj 33.2 34.4 33.8
Dagbani dag 32.4 33.9 33.1
Kakwa keo 34.3 30.3 32.3
Berber ber 24.2 38.6 31.4
Ma'di mhi 25.8 27.9 26.9
Fulani ful 21.3 32.2 26.8
Kwamba rwm 26.9 25.9 26.4
Aringa luc 22.8 29.1 26.0
Luhya luy 17.7 26.7 22.2
Lendu led 18.4 24.5 21.4
Dinka din 17.4 25.5 21.4
Dagaare dga 16.0 25.4 20.7
Kupsabiny kpz 17.2 22.7 19.9
Pokot pok 17.1 20.3 18.7
Ik ikx 15.7 20.9 18.3
Kalenjin kln 13.0 22.6 17.8
Kanuri kau 9.8 22.2 16.0
Ikposo kpo 4.0 18.8 11.4

Comprehension, understanding and reasoning

AfroBench-Lite, 15 languages (Belebele 12), 5-shot, 100 documents per task. Accuracy, or exact match for AfriMGSM. Best score per task in bold.

Task Sunflower-27B Sunflower-9B AfriqueQwen3.5-9B Instruct v1 AfriqueQwen3.5-4B Instruct v1 Qwen3.8-27B Qwen3.5-27B Qwen3.5-9B
Reasoning and comprehension
AfriMMLU 62.8 47.8 54.5 44.4 58.9 62.9 42.1
AfriMGSM 40.3 16.7 24.8 34.9 32.2 35.7 21.3
Belebele 82.2 68.2 69.2 62.4 72.0 75.2 58.8
Language understanding
AfriXNLI 69.9 61.5 71.6 68.1 66.9 66.3 57.8
SIB-200 86.5 83.5 84.7 77.5 82.4 82.8 72.4
Injongo intent 87.7 81.9 80.5 67.2 83.7 85.1 71.2

Sunflower-27B leads four of six tasks and is within 0.1 of the best on AfriMMLU. The 9B is tuned for translation; for comprehension and maths, use the 27B.

Adaptive thinking. With thinking enabled the model reasons where it helps and answers directly elsewhere, so maths and knowledge improve and everything else holds (generative AfroBench-Lite tasks, Sunflower-27B):

Task Thinking off Thinking on
AfriMGSM 40.3 60.5
AfriMMLU 66.6 71.6
Belebele 84.6 84.3
SIB-200 88.0 87.3
Translation (chrF, 3 languages) 64.3 64.4

SAHARA

Scores from the SAHARA evaluation server, 18 tasks in four clusters. Sunflower-27B is the top open-weight model on the board and sixth overall, behind five closed models.

Cluster Sunflower-27B, adaptive Sunflower-27B, thinking off Sunflower-9B, adaptive Qwen3.8-27B Qwen3.5-27B
Knowledge, comprehension and reasoning 64.8 56.5 43.5 69.4 54.6
Text classification 45.2 45.2 39.5 42.4 43.6
Text generation 17.7 17.7 16.1 11.3 11.7
Token-level tasks 40.9 41.4 13.6 45.4 33.8
SAHARA score 42.15 40.22 28.19 42.12 35.95

Training for African languages lifts text generation by 6.4 points and classification by 2.8 over the starting checkpoint, with adaptive thinking closing most of the knowledge gap.

Usage

The model follows the Qwen3.8 chat template. Leave thinking enabled (the default) for general use; the model decides per prompt whether to reason. For bulk translation, set enable_thinking=False: translation quality is the same and generation is faster. Use greedy decoding for translation and a temperature around 0.7 for open-ended generation.

Prompting

Translation:

Translate to {language}: {text}

{language} is the English name of the target language (Luganda, Acholi, Swahili). The source language is inferred; naming it (Translate from Luganda to English: ...) helps with short or ambiguous input.

For other tasks, write the instruction in English and state the reply language (... Reply in Yoruba.), or write the instruction in the target language: the model answers in the language it is addressed in. Multi-turn chat works as normal.

vLLM

import os

os.environ["VLLM_ENABLE_V1_MULTIPROCESSING"] = "0"

from transformers import AutoProcessor
from vllm import LLM, SamplingParams

MODEL_ID = "Sunbird/Sunflower-Qwen3.8-27B"

llm = LLM(
    MODEL_ID,
    quantization="fp8",  # ~55 GB of weights -> ~28 GB
    dtype="bfloat16",
    max_model_len=2048,
    gpu_memory_utilization=0.9,
    limit_mm_per_prompt={"image": 0, "video": 0},  # text only
)
tokenizer = AutoProcessor.from_pretrained(MODEL_ID).tokenizer


def make_prompt(text: str, target_language: str) -> str:
    """Build a translation prompt with thinking disabled."""
    return tokenizer.apply_chat_template(
        [{"role": "user", "content": f"Translate to {target_language}: {text}"}],
        add_generation_prompt=True,
        enable_thinking=False,
        tokenize=False,
    )


examples = [
    ("Good morning, how are you?", "Luganda"),
    ("The children are going to school today.", "Acholi"),
    ("I would like to buy three kilograms of rice.", "Swahili"),
]
outputs = llm.generate(
    [make_prompt(text, lang) for text, lang in examples],
    SamplingParams(temperature=0.0, max_tokens=256),
)
for (text, lang), out in zip(examples, outputs):
    print(f"[{lang}] {text}\n  -> {out.outputs[0].text.strip()}")

Transformers

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

MODEL_ID = "Sunbird/Sunflower-Qwen3.8-27B"

model = AutoModelForImageTextToText.from_pretrained(
    MODEL_ID, dtype=torch.bfloat16, device_map="auto"
).eval()
tokenizer = AutoProcessor.from_pretrained(MODEL_ID).tokenizer

messages = [
    {"role": "user", "content": "Translate to Luganda: The children are going to school today."}
]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=False,
    return_tensors="pt",
    return_dict=True,
).to(model.device)
output = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(
    tokenizer.decode(
        output[0][inputs["input_ids"].shape[1] :], skip_special_tokens=True
    ).strip()
)

With thinking enabled the reply starts with a <think>...</think> block; allow max_new_tokens of 2048 or more for reasoning tasks.

Training

Starting from Qwen3.8-27B: continued pretraining on African-language text; supervised fine-tuning for translation, comprehension, tagging and instruction following; then reinforcement learning (GRPO) with rewards for translation quality, factuality, language consistency and instruction following. Adaptive reasoning was trained so the model chooses when to think.

Limitations

  • Coverage is uneven. Quality tracks training data; see the tiers above, and check basic-tier output with a speaker before use.
  • Translation is the most optimised path. General chat and reasoning work well, but the starting checkpoint keeps an edge on general knowledge and NER (see SAHARA).
  • Fluent is not the same as correct. Like any LLM it can produce confident, wrong output, and errors are harder to spot in low-resource languages.

License

Apache 2.0.

Citation

@misc{sunflower_qwen38_27b,
  title        = {Sunflower-Qwen3.8-27B},
  author       = {Sunbird AI},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Sunbird/Sunflower-Qwen3.8-27B}}
}
Downloads last month
553
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sunbird/Sunflower-Qwen3.8-27B

Base model

Qwen/Qwen3.8-27B
Finetuned
(472)
this model

Collection including Sunbird/Sunflower-Qwen3.8-27B