BlueMagpie-TTS Usage
This is an inference checkpoint. Install the local package first:
git clone https://github.com/voidful/BlueMagpie-TTS
cd BlueMagpie-TTS
pip install -e ".[speaker]"
pip install soundfile
Download and load:
from huggingface_hub import snapshot_download
from bluemagpie import BlueMagpieModel, speaker_embedding_from_wav
model_dir = snapshot_download("OpenFormosa/BlueMagpie-TTS")
model = BlueMagpieModel.from_local(model_dir, training=False, device="cuda")
Generate a short utterance:
import soundfile as sf
reference_embedding = speaker_embedding_from_wav(
"reference_speaker.wav", window_sec=3.0, hop_sec=1.5
)
wav = model.generate(
target_text="這是合成語音測試。",
speaker_centroid=reference_embedding,
cfg_value=2.0,
inference_timesteps=10,
retry_badcase=True,
retry_badcase_ratio_threshold=6.0,
stop_threshold=0.65,
stop_consecutive=2,
)
sf.write("short.wav", wav.detach().cpu().numpy(), model.sample_rate)
Generation modes:
# centroid-only
wav = model.generate(
target_text="這是指定 speaker centroid 的測試。",
speaker_centroid=centroid_tensor,
cfg_value=2.0,
inference_timesteps=10,
stop_threshold=0.65,
stop_consecutive=2,
)
# speaker-reference-audio-only; extract once, no reference transcript required
reference_embedding = speaker_embedding_from_wav(
"reference_speaker.wav", window_sec=3.0, hop_sec=1.5
)
wav = model.generate(
target_text="今天的 meeting 依照原定時間進行。",
speaker_centroid=reference_embedding,
cfg_value=2.0,
inference_timesteps=10,
stop_threshold=0.65,
stop_consecutive=2,
)
Centroid demo:
import torch
table = torch.load(model_dir + "/checkpoints/speaker_centroids.pt", map_location="cpu")
wav = model.generate(
target_text="這是 centroid demo 的測試。",
speaker_centroid=table["centroids"][0],
cfg_value=2.0,
inference_timesteps=10,
)
Streaming:
import torch
import soundfile as sf
chunks = []
for chunk in model.generate_streaming(
target_text="這是一段串流語音合成測試。",
speaker_centroid=reference_embedding,
cfg_value=2.0,
inference_timesteps=10,
):
chunks.append(chunk.detach().cpu())
wav = torch.cat(chunks, dim=-1)
sf.write("streaming.wav", wav.numpy(), model.sample_rate)
Long text:
python scripts/generate_tts.py \
--checkpoint /path/to/model-snapshot \
--mode speaker-reference-audio-only \
--speaker-reference-wav /path/to/rights-cleared-reference.wav \
--reference-audio-conditioning speaker-embedding \
--text-file long_text.txt \
--chunk-chars 80 \
--min-chunk-chars 12 \
--target-chars-per-sec 4.0 \
--stop-threshold 0.65 \
--stop-consecutive 2 \
--crossfade-ms 80 \
--chunk-rms-match-db 4 \
--chunk-edge-fade-ms 80 \
--continuation-context-sec 0 \
--no-retry-badcase \
--out long.wav
Quality-first short or medium text:
python scripts/generate_tts_quality.py \
--checkpoint /path/to/model-snapshot \
--mode speaker-reference-audio-only \
--speaker-reference-wav /path/to/rights-cleared-reference.wav \
--text "這是離線品質優先的語音合成測試。" \
--candidates 10 \
--asr-backend whisper \
--asr-model /path/to/compatible-asr-checkpoint \
--candidate-speaker-weight 0.05 \
--candidate-boundary-speaker-weight 0.1 \
--candidate-max-boundary-speaker-drop 0.03 \
--target-chars-per-sec 4.2 \
--out quality.wav
This opt-in path is approximately proportional to the candidate count in
latency. It fails closed if no candidate passes the speaker-boundary gate.
--allow-boundary-fallback is a research override, not part of the measured
contract. The quality policy is promoted only for speaker-reference
short/medium text; keep using generate_tts.py for centroid and long-form
production requests.
This path extracts the windowed speaker embedding once and reuses it for every
chunk. Each request gets a new random seed that is reused inside that request;
the release gate varies request seeds, so stability does not depend on one
global fixed seed. --reference-audio-conditioning tokens enables the
raw-reference-token research backend; it is not the production default.
Recommended generation defaults:
cfg_value=2.0inference_timesteps=10retry_badcase=Falsefor the measured long-form contractretry_badcase_ratio_threshold=6.0so retry detection is not a pace capstop_threshold=0.65,stop_consecutive=2for offline generationmax_len=2000unless a longer chunk is intentionally neededtarget-chars-per-sec=4.0for controlled long-form chunkingcontinuation-context-sec=0until generated-context continuation passes its own gate- Reference wavs used to extract speaker embeddings should be at least 3 seconds.
The hosted interactive demo uses a separate endpoint-only policy with the same weights: a 0.50 stop threshold that relaxes after 75% of the native-rate duration estimate to 0.05 at 95%, one stop hit, and a native-rate hard cap plus one latent step. It applies pace correction after generation. The offline defaults above remain the fixed evaluation contract.
Safety:
- Use only rights-cleared reference audio or speaker embeddings.
- Do not present generated speech as a real person or real notification unless that is explicitly authorized and reviewed.