Instructions to use Masterx/magpie-tts-multilingual-357m-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use Masterx/magpie-tts-multilingual-357m-ONNX with NeMo:
# tag did not correspond to a valid NeMo domain.
- Notebooks
- Google Colab
- Kaggle
Magpie-TTS Multilingual 357M — ONNX
ONNX export of NVIDIA's Magpie-TTS Multilingual 357M (checkpoint v2607) and its NanoCodec 22.05 kHz decoder. It was made for WinSTT, which runs it on CPU through ONNX Runtime with a native Rust decode loop. The graphs are plain ONNX and work with any ORT runtime.
All credit for the model goes to NVIDIA. This repo only re-packages the weights.
License
Licensed by NVIDIA Corporation under the NVIDIA Open Model License. See LICENSE (the full agreement)
and NOTICE. Any redistribution must keep both files.
What is in the repo
| File | Bytes | Notes |
|---|---|---|
text_encoder.onnx |
439,566,205 | fp32 text encoder |
decoder_step.onnx |
455,444,206 | fp32 decoder step (prefill and single-frame) |
local_step.onnx |
156,330,855 | fp32 local transformer step |
text_encoder_int8.onnx |
390,242,041 | int8 text encoder |
decoder_step_int8.onnx |
119,720,820 | int8 decoder step |
local_step_int8.onnx |
113,947,162 | int8 local transformer step |
codec_decoder.onnx |
128,130,025 | NanoCodec decoder, fp32 (shared by both sets) |
audio_embeddings.bin |
99,483,648 | codebook embeddings, f32 [16, 2024, 768] |
speaker_context.bin |
3,333,120 | baked speaker contexts, f32 [5, 217, 768] |
tokenizer/* |
20,600,373 | NeMo tokenizer config + G2P dictionaries (8 files) |
LICENSE |
10,494 | NVIDIA Open Model License Agreement |
NOTICE |
67 | required attribution notice |
Download size: int8 set 875,467,750 bytes, fp32 set 1,302,898,993 bytes (both include the shared files).
The *_int8.onnx files are dynamic int8 versions of the three transformer graphs: MatMul/Gemm weights,
per-channel, made with onnxruntime.quantization.quantize_dynamic. The codec is fp32 in both sets. A full
fp32 set is the three plain graphs plus the shared files. The int8 set is the three _int8 graphs plus
the same shared files.
Languages and voices
- Languages: English, German, Spanish, French, Italian, Brazilian Portuguese, Hindi, Arabic, Korean and Vietnamese. The upstream checkpoint also covers Mandarin and Japanese. Those two need jieba or OpenJTalk segmenters, so they are not supported by WinSTT's tokenizer port. The graphs themselves do not depend on the language.
- Voices: five baked speakers, stored in
speaker_context.binin this order: Aria (F), Jason (M), John (M), Leo (M), Sofia (F). Every voice speaks every language. Upstream removed zero-shot cloning, so the export has none.
Graphs
The model is split into four graphs plus two host-side tables. The NeMo inference loop runs on the host: CFG pairing, the attention prior, sampling and EOS detection.
| Graph | Inputs | Outputs |
|---|---|---|
text_encoder |
text int64 [1, T] |
cross_k, cross_v f32 [12, 2, T, 128] (per-layer cross-attention K/V, duplicated for the CFG pair) |
decoder_step |
x f32 [B, T, 768], past_k/past_v f32 [12, B, 12, Tp, 64], cross_k/cross_v f32 [12, B, Tt, 128], cond_mask f32 [B, Tt], prior f32 [B, Tt] |
logits f32 [B, 32384] (16 × 2024), dec_out f32 [B, 768], align f32 [B, Tt], new_k/new_v f32 [12, B, 12, Tp+T, 64] |
local_step |
x_in f32 [B, 768], cb int64 scalar (flat codebook index 0..15), past_k/past_v f32 [2, B, 12, Tp, 64] |
logits f32 [B, 2024], new_k/new_v |
codec_decoder |
codes int64 [1, 8, T] |
audio f32 [1, T × 1024] at 22,050 Hz |
Host tables (raw little-endian f32):
audio_embeddings.bin: [16, 2024, 768]. The codebook embeddings for 8 codebooks × frame-stacking 2.speaker_context.bin: [5, 217, 768]. The baked speaker contexts.
Tokenizer: tokenizer/ holds the NeMo tokenizer configs and pronunciation dictionaries (IPA
G2P for en/de/es/pt-br/hi, plus character tokenizers). The text ids must end with EOS 3358.
Decode loop (NeMo generate_speech defaults):
- Prefill
[speaker context ; BOS frame]for the conditional row and[zeros ; BOS frame]for the unconditional row. The unconditional row attends only text position 0 and gets no prior. - Mix logits with CFG 2.5.
- Apply the attention prior: eps 0.1, lookahead 6.
- The local transformer samples 16 codes per step with temperature 0.6 and top-k 80.
- Stop at the first EOS (2017) in either the sampled or the argmax frame. BOS is 2016.
The steps cap at 250 frame pairs (about 23 s), so split long text.
Export method
- Software: NeMo 3.0.0 (
MagpieTTSModel) and PyTorch 2.14.1. Exported withtorch.onnx.export(TorchScript exporter) at opset 18. The decoder and local transformer are rewritten as single-step graphs with explicit KV caches, using NeMo's own cache semantics. - Codec: the NanoCodec
CausalHiFiGANDecoderis exported with weight-norm parametrizations folded into plain conv weights. The graph is lighter and gives the same output.
Validation
| Check | Result |
|---|---|
| Graph vs NeMo modules (max abs) | text encoder ≤ 2.1e-5, decoder step ≤ 3.0e-5 (prefill, cached steps, 250-frame cache), local step ≤ 2.7e-5, codec ≤ 4.5e-5 (folded codec vs NeMo 1.9e-5) |
| End-to-end greedy codes, NeMo vs ONNX | 12/12 sentences (en, de, fr, hi × 3): identical codes and frame counts (token agreement 1.0) |
| Native Rust loop vs Python ORT reference (greedy codes) | 12/12 sentences identical (token agreement 1.0) |
| WER/CER (whisper-small, sampled) | see below |
Intelligibility: whisper-small transcriptions of sampled (NeMo-default) generations, 3–10 sentences per language, across all five voices. Text is normalized with the whisper normalizers.
| Language | Sentences | fp32 WER / CER | int8 WER / CER |
|---|---|---|---|
| English | 10 | 0.0% / 0.0% | 0.0% / 0.0% |
| German | 5 | 0.0% / 0.0% | 0.0% / 0.0% |
| Spanish | 5 | 2.6% / 2.0% | 2.6% / 2.0% |
| French | 5 | 2.6% / 1.4% | 2.6% / 1.4% |
| Italian | 5 | 0.0% / 0.0% | 0.0% / 0.0% |
| Portuguese (BR) | 5 | 2.6% / 1.5% | 2.6% / 1.5% |
| Hindi | 5 | 22.0% / 11.4% | 22.0% / 9.6% |
| Arabic | 3 | 6.7% / 2.7% | 13.3% / 5.4% |
| Korean | 3 | 0.0% / 0.0% | 0.0% / 0.0% |
| Vietnamese | 3 | 14.3% / 4.4% | 4.8% / 1.1% |
How these were produced:
- The fp32 column comes from WinSTT's Rust engine. The int8 column comes from the Python ORT reference loop. The two loops use different random generators, so their sampled outputs differ.
- The sentences are short, so one word moves the 3-sentence sets by about 5–7 points.
- Most Hindi errors are whisper-small spellings (for example मौसम → मुसम, आरक्षित → आरक्षिट) and digits ("दो" → "2"). One fp32 sample is transcribed with a repeated word.
Speed: Measured on CPU only (24 logical threads), on a shared machine. The best quiet-machine figure was about 6× real time for fp32: roughly 4× for generation, plus about 2× for the NanoCodec decoder. Under load the measurements were 11× (int8, Python ORT) and 19× (fp32, unoptimized Rust test build). Expect it to be slower than real time on CPU. Most of the time goes to the codec and to the 16 local-transformer steps per frame pair.
Tip for ONNX Runtime users: disable intra-op spinning (session.intra_op.allow_spinning = 0) when you run
the four sessions in one process. With spinning on, the idle thread pools fight the active one. That made
generation about 3× slower in our tests.
Model tree for Masterx/magpie-tts-multilingual-357m-ONNX
Base model
nvidia/magpie_tts_multilingual_357m