Qwen3 TTS Tiny Turkish

A small Turkish text-to-speech model with the structure of Qwen3-0.6B-Base: 199M parameters, 4.6× smaller. On Freya-TR-Eval (495 sentences) it reaches 1.28% WER and 0.42% CER, the same level as Qwen3 0.6B Turkish (1.4% / 0.29%).

Give it a clean 5–10 second Turkish reference recording to choose the voice. It does not require per-voice fine-tuning or a fixed voice list.

Samples

Reference voices were not used in training.

Voice Text Audio
1 İstanbul'da sabahları vapurla karşıya geçmek, şehrin en eski alışkanlıklarından biri. Martıların sesi, uzaktan gelen düdükler ve suyun üstünde açılan yol, güne sakin bir başlangıç yapar.
2 Anadolu'nun küçük kasabalarında akşam olunca sokaklar birden sessizleşir. Kahvede tavla oynayanlar evlerine dağılır, çeşmeden su dolduran çocuklar koşarak uzaklaşır. Gün, yavaşça karanlığa bırakır yerini.
3 Bilim insanları yıllardır uzayın derinliklerinden gelen sinyalleri inceliyor. Bu sinyallerin bir kısmı yıldızların doğumuna, bir kısmı ise çok uzak galaksilerin hareketine işaret ediyor. Her yeni ölçüm, evrenin bilinen sınırlarını biraz daha genişletiyor.
4 Bir kitabı yeniden okumak, aynı metinden farklı bir anlam çıkarmak gibidir. Yıllar geçtikçe cümleler değişmez, ama onları okuyan kişi artık aynı kişi değildir. Bu yüzden bazı kitaplar, her okunuşta başka bir kitaba dönüşür.

Results

120 Freya-TR-Eval sentences, each read with a real speaker's voice the models never heard in training. WER/CER from Whisper large-v3, speaker similarity from ECAPA, naturalness from UTMOS22.

Model WER CER Speaker similarity UTMOS
Qwen Tiny SFT 3.2% 0.8% 0.45 3.41
Qwen Tiny RL (this model) 3.3% 0.7% 0.46 3.58
Qwen3 0.6B Turkish 2.0% 0.5% 0.39 3.76
  • Speaker similarity is good: the model follows a new voice more closely than Qwen3 0.6B Turkish (0.46 vs 0.39 with real speakers; 0.71 vs 0.45 with held-out synthetic voices).
  • Accuracy: Whisper writes some spoken numbers as digits, which counts as an error for every model. Scoring numbers as words lowers WER by about 0.3 points.
  • Naturalness: Qwen3 0.6B Turkish still sounds more natural; RL closed part of the gap.
  • Scores were measured with Qwen's original sampling settings (0.9 / 0.9 / 1.05); this repository's defaults sound smoother.
Qwen Tiny Qwen3 0.6B Turkish
Parameters (without codec) 199M 915M
GPU, one sentence per call (L4) RTF 1.08, 0.84 GB RTF 2.19, 2.18 GB

It also runs on a MacBook CPU (Apple M5, 4 threads, PyTorch FP32) faster than real time, with under 4 GB of RAM.

Model

  • Base model: Qwen3-0.6B-Base at revision 5d83992
  • Parameters: ~370M packaged (199M talker and speaker encoder, plus the 171M speech tokenizer)
  • Talker: 8 layers × 512 (base: 28 × 1024); code predictor 2 layers × 1024 (base: 5 × 1024)
  • Audio frame rate: 12 Hz
  • Sample rate: 24 kHz
  • Language: Turkish (Auto conditioning; no new language token)
  • Voice control: reference recording → speaker x-vector
  • Output: non-streaming

The layers and channels were picked from the base weights by how much they matter on Turkish speech, then trained. Only the talker was trained. The speech tokenizer is unchanged; the speaker encoder is the base encoder with its output cut to 512 channels.

Training

Data: Alania Turkish Synthetic Speech, cc-by config at revision 30639fa.

  • Training clips: 1,451,721, up to 12 seconds each
  • Audio: ~1,863 hours
  • Voices: 2,728 synthetic voices
  • Held-out evaluation: 24 voices, 12,834 clips

Text was kept in its spoken form.

  1. SFT: 90,000 steps of 64 clips.
  2. RL: group-relative policy optimisation on the whole talker, 253 steps of 16 sentences × 8 samples. Rewards: speaker similarity to the reference (WavLM) and naturalness (UTMOS22), with a KL penalty to the SFT weights. RL made speech a little more natural; the difference from SFT is small.

Compared with Qwen3 0.6B Turkish. That model is the full Qwen3-TTS 0.6B Base, fine-tuned for one epoch on about 2,277 hours of real read speech from 54 speakers. This model is 4.6× smaller and learned from synthetic speech with many more voices, which is why Qwen3 0.6B Turkish sounds more natural and this model follows new voices more closely.

Usage

import soundfile as sf
import torch
from qwen_tts import Qwen3TTSModel

model = Qwen3TTSModel.from_pretrained("erkamk/qwen3-tts-tiny-tr", device_map="cpu", dtype=torch.float32)
voice = model.create_voice_clone_prompt(ref_audio="reference.wav", x_vector_only_mode=True)

text = "Merhaba, size nasıl yardımcı olabilirim?"
text = text.replace("İ", "i").replace("I", "ı")  # the training text had no capital İ or I
wavs, rate = model.generate_voice_clone(
    text=text, language="Auto", voice_clone_prompt=voice, non_streaming_mode=True,
    max_new_tokens=25 + 2 * len(text),
    temperature=0.7, subtalker_temperature=0.6, repetition_penalty=1.0,
)
sf.write("out.wav", wavs[0], rate)
  • Use a clean 5–10 second reference with one speaker.
  • Recommended settings are the ones above and are also this repository's defaults. Qwen's original defaults (0.9 / 0.9 / 1.05) can cause small breaks with real-voice references.
  • For paragraphs, generate sentence by sentence, or pass a list of sentences to batch them.
  • Write numbers, prices and dates as words, and write capital İ and I as i and ı (the line above does this).

Known issues and next steps

  • Speech became slightly slower after RL.
  • With real-voice references, the library's default sampling can cause small breaks; the recommended settings, now this repository's defaults, reduce them.
  • Very long input can end early and drop the last words: training clips were at most 12 seconds, and our longest test paragraph (about 18 seconds of speech) came out 3–6 seconds short in a single call.

More RL training is planned to fix these. If you notice a problem, please open a discussion on this repository.

License and attribution

Weights: Apache 2.0, inherited from Qwen3-TTS.

Alania Turkish Synthetic Speech by PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-synthetic-speech-tr), licensed under CC BY 4.0.

Only the cc-by config was used, for training and for the sample reference voices. Its FineWeb-2 text is used under ODC-By 1.0 and remains subject to Common Crawl's terms of use.

@misc{patientdesk2026alaniasyntheticspeech,
  title        = {Alania Turkish Synthetic Speech},
  author       = {{PatientDesk AI}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/cloud0day3/alania-synthetic-speech-tr}},
  note         = {3,248 h of AI-generated Turkish speech, CC BY 4.0 / CC BY-SA 4.0}
}

You are responsible for using voice cloning lawfully and with consent, including local rules on synthetic voices and disclosure.

Downloads last month
112
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for erkamk/qwen3-tts-tiny-tr

Finetuned
(31)
this model

Dataset used to train erkamk/qwen3-tts-tiny-tr

Space using erkamk/qwen3-tts-tiny-tr 1