Text-to-Speech
ONNX
NeMo
tts
magpie
nanocodec
winstt

Magpie-TTS Multilingual 357M — ONNX

ONNX export of NVIDIA's Magpie-TTS Multilingual 357M (checkpoint v2607) and its NanoCodec 22.05 kHz decoder. It was made for WinSTT, which runs it on CPU through ONNX Runtime with a native Rust decode loop. The graphs are plain ONNX and work with any ORT runtime.

All credit for the model goes to NVIDIA. This repo only re-packages the weights.

License

Licensed by NVIDIA Corporation under the NVIDIA Open Model License. See LICENSE (the full agreement) and NOTICE. Any redistribution must keep both files.

What is in the repo

File Bytes Notes
text_encoder.onnx 439,566,205 fp32 text encoder
decoder_step.onnx 455,444,206 fp32 decoder step (prefill and single-frame)
local_step.onnx 156,330,855 fp32 local transformer step
text_encoder_int8.onnx 390,242,041 int8 text encoder
decoder_step_int8.onnx 119,720,820 int8 decoder step
local_step_int8.onnx 113,947,162 int8 local transformer step
codec_decoder.onnx 128,130,025 NanoCodec decoder, fp32 (shared by both sets)
audio_embeddings.bin 99,483,648 codebook embeddings, f32 [16, 2024, 768]
speaker_context.bin 3,333,120 baked speaker contexts, f32 [5, 217, 768]
tokenizer/* 20,600,373 NeMo tokenizer config + G2P dictionaries (8 files)
LICENSE 10,494 NVIDIA Open Model License Agreement
NOTICE 67 required attribution notice

Download size: int8 set 875,467,750 bytes, fp32 set 1,302,898,993 bytes (both include the shared files).

The *_int8.onnx files are dynamic int8 versions of the three transformer graphs: MatMul/Gemm weights, per-channel, made with onnxruntime.quantization.quantize_dynamic. The codec is fp32 in both sets. A full fp32 set is the three plain graphs plus the shared files. The int8 set is the three _int8 graphs plus the same shared files.

Languages and voices

  • Languages: English, German, Spanish, French, Italian, Brazilian Portuguese, Hindi, Arabic, Korean and Vietnamese. The upstream checkpoint also covers Mandarin and Japanese. Those two need jieba or OpenJTalk segmenters, so they are not supported by WinSTT's tokenizer port. The graphs themselves do not depend on the language.
  • Voices: five baked speakers, stored in speaker_context.bin in this order: Aria (F), Jason (M), John (M), Leo (M), Sofia (F). Every voice speaks every language. Upstream removed zero-shot cloning, so the export has none.

Graphs

The model is split into four graphs plus two host-side tables. The NeMo inference loop runs on the host: CFG pairing, the attention prior, sampling and EOS detection.

Graph Inputs Outputs
text_encoder text int64 [1, T] cross_k, cross_v f32 [12, 2, T, 128] (per-layer cross-attention K/V, duplicated for the CFG pair)
decoder_step x f32 [B, T, 768], past_k/past_v f32 [12, B, 12, Tp, 64], cross_k/cross_v f32 [12, B, Tt, 128], cond_mask f32 [B, Tt], prior f32 [B, Tt] logits f32 [B, 32384] (16 × 2024), dec_out f32 [B, 768], align f32 [B, Tt], new_k/new_v f32 [12, B, 12, Tp+T, 64]
local_step x_in f32 [B, 768], cb int64 scalar (flat codebook index 0..15), past_k/past_v f32 [2, B, 12, Tp, 64] logits f32 [B, 2024], new_k/new_v
codec_decoder codes int64 [1, 8, T] audio f32 [1, T × 1024] at 22,050 Hz

Host tables (raw little-endian f32):

  • audio_embeddings.bin: [16, 2024, 768]. The codebook embeddings for 8 codebooks × frame-stacking 2.
  • speaker_context.bin: [5, 217, 768]. The baked speaker contexts.

Tokenizer: tokenizer/ holds the NeMo tokenizer configs and pronunciation dictionaries (IPA G2P for en/de/es/pt-br/hi, plus character tokenizers). The text ids must end with EOS 3358.

Decode loop (NeMo generate_speech defaults):

  1. Prefill [speaker context ; BOS frame] for the conditional row and [zeros ; BOS frame] for the unconditional row. The unconditional row attends only text position 0 and gets no prior.
  2. Mix logits with CFG 2.5.
  3. Apply the attention prior: eps 0.1, lookahead 6.
  4. The local transformer samples 16 codes per step with temperature 0.6 and top-k 80.
  5. Stop at the first EOS (2017) in either the sampled or the argmax frame. BOS is 2016.

The steps cap at 250 frame pairs (about 23 s), so split long text.

Export method

  • Software: NeMo 3.0.0 (MagpieTTSModel) and PyTorch 2.14.1. Exported with torch.onnx.export (TorchScript exporter) at opset 18. The decoder and local transformer are rewritten as single-step graphs with explicit KV caches, using NeMo's own cache semantics.
  • Codec: the NanoCodec CausalHiFiGANDecoder is exported with weight-norm parametrizations folded into plain conv weights. The graph is lighter and gives the same output.

Validation

Check Result
Graph vs NeMo modules (max abs) text encoder ≤ 2.1e-5, decoder step ≤ 3.0e-5 (prefill, cached steps, 250-frame cache), local step ≤ 2.7e-5, codec ≤ 4.5e-5 (folded codec vs NeMo 1.9e-5)
End-to-end greedy codes, NeMo vs ONNX 12/12 sentences (en, de, fr, hi × 3): identical codes and frame counts (token agreement 1.0)
Native Rust loop vs Python ORT reference (greedy codes) 12/12 sentences identical (token agreement 1.0)
WER/CER (whisper-small, sampled) see below

Intelligibility: whisper-small transcriptions of sampled (NeMo-default) generations, 3–10 sentences per language, across all five voices. Text is normalized with the whisper normalizers.

Language Sentences fp32 WER / CER int8 WER / CER
English 10 0.0% / 0.0% 0.0% / 0.0%
German 5 0.0% / 0.0% 0.0% / 0.0%
Spanish 5 2.6% / 2.0% 2.6% / 2.0%
French 5 2.6% / 1.4% 2.6% / 1.4%
Italian 5 0.0% / 0.0% 0.0% / 0.0%
Portuguese (BR) 5 2.6% / 1.5% 2.6% / 1.5%
Hindi 5 22.0% / 11.4% 22.0% / 9.6%
Arabic 3 6.7% / 2.7% 13.3% / 5.4%
Korean 3 0.0% / 0.0% 0.0% / 0.0%
Vietnamese 3 14.3% / 4.4% 4.8% / 1.1%

How these were produced:

  • The fp32 column comes from WinSTT's Rust engine. The int8 column comes from the Python ORT reference loop. The two loops use different random generators, so their sampled outputs differ.
  • The sentences are short, so one word moves the 3-sentence sets by about 5–7 points.
  • Most Hindi errors are whisper-small spellings (for example मौसम → मुसम, आरक्षित → आरक्षिट) and digits ("दो" → "2"). One fp32 sample is transcribed with a repeated word.

Speed: Measured on CPU only (24 logical threads), on a shared machine. The best quiet-machine figure was about 6× real time for fp32: roughly 4× for generation, plus about 2× for the NanoCodec decoder. Under load the measurements were 11× (int8, Python ORT) and 19× (fp32, unoptimized Rust test build). Expect it to be slower than real time on CPU. Most of the time goes to the codec and to the 16 local-transformer steps per frame pair.

Tip for ONNX Runtime users: disable intra-op spinning (session.intra_op.allow_spinning = 0) when you run the four sessions in one process. With spinning on, the idle thread pools fight the active one. That made generation about 3× slower in our tests.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Masterx/magpie-tts-multilingual-357m-ONNX

Quantized
(6)
this model