BlueMagpie-TTS / USAGE.md
voidful's picture
Document final endpoint policy and rejected speaker-tail pilot
4e11051 verified
|
Raw
History Blame Contribute Delete
5.46 kB

BlueMagpie-TTS Usage

This is an inference checkpoint. Install the local package first:

git clone https://github.com/voidful/BlueMagpie-TTS
cd BlueMagpie-TTS
pip install -e ".[speaker]"
pip install soundfile

Download and load:

from huggingface_hub import snapshot_download
from bluemagpie import BlueMagpieModel, speaker_embedding_from_wav

model_dir = snapshot_download("OpenFormosa/BlueMagpie-TTS")
model = BlueMagpieModel.from_local(model_dir, training=False, device="cuda")

Generate a short utterance:

import soundfile as sf

reference_embedding = speaker_embedding_from_wav(
    "reference_speaker.wav", window_sec=3.0, hop_sec=1.5
)
wav = model.generate(
    target_text="這是合成語音測試。",
    speaker_centroid=reference_embedding,
    cfg_value=2.0,
    inference_timesteps=10,
    retry_badcase=True,
    retry_badcase_ratio_threshold=6.0,
    stop_threshold=0.65,
    stop_consecutive=2,
)
sf.write("short.wav", wav.detach().cpu().numpy(), model.sample_rate)

Generation modes:

# centroid-only
wav = model.generate(
    target_text="這是指定 speaker centroid 的測試。",
    speaker_centroid=centroid_tensor,
    cfg_value=2.0,
    inference_timesteps=10,
    stop_threshold=0.65,
    stop_consecutive=2,
)

# speaker-reference-audio-only; extract once, no reference transcript required
reference_embedding = speaker_embedding_from_wav(
    "reference_speaker.wav", window_sec=3.0, hop_sec=1.5
)
wav = model.generate(
    target_text="今天的 meeting 依照原定時間進行。",
    speaker_centroid=reference_embedding,
    cfg_value=2.0,
    inference_timesteps=10,
    stop_threshold=0.65,
    stop_consecutive=2,
)

Centroid demo:

import torch

table = torch.load(model_dir + "/checkpoints/speaker_centroids.pt", map_location="cpu")
wav = model.generate(
    target_text="這是 centroid demo 的測試。",
    speaker_centroid=table["centroids"][0],
    cfg_value=2.0,
    inference_timesteps=10,
)

Streaming:

import torch
import soundfile as sf

chunks = []
for chunk in model.generate_streaming(
    target_text="這是一段串流語音合成測試。",
    speaker_centroid=reference_embedding,
    cfg_value=2.0,
    inference_timesteps=10,
):
    chunks.append(chunk.detach().cpu())

wav = torch.cat(chunks, dim=-1)
sf.write("streaming.wav", wav.numpy(), model.sample_rate)

Long text:

python scripts/generate_tts.py \
  --checkpoint /path/to/model-snapshot \
  --mode speaker-reference-audio-only \
  --speaker-reference-wav /path/to/rights-cleared-reference.wav \
  --reference-audio-conditioning speaker-embedding \
  --text-file long_text.txt \
  --chunk-chars 80 \
  --min-chunk-chars 12 \
  --target-chars-per-sec 4.0 \
  --stop-threshold 0.65 \
  --stop-consecutive 2 \
  --crossfade-ms 80 \
  --chunk-rms-match-db 4 \
  --chunk-edge-fade-ms 80 \
  --continuation-context-sec 0 \
  --no-retry-badcase \
  --out long.wav

Quality-first short or medium text:

python scripts/generate_tts_quality.py \
  --checkpoint /path/to/model-snapshot \
  --mode speaker-reference-audio-only \
  --speaker-reference-wav /path/to/rights-cleared-reference.wav \
  --text "這是離線品質優先的語音合成測試。" \
  --candidates 10 \
  --asr-backend whisper \
  --asr-model /path/to/compatible-asr-checkpoint \
  --candidate-speaker-weight 0.05 \
  --candidate-boundary-speaker-weight 0.1 \
  --candidate-max-boundary-speaker-drop 0.03 \
  --target-chars-per-sec 4.2 \
  --out quality.wav

This opt-in path is approximately proportional to the candidate count in latency. It fails closed if no candidate passes the speaker-boundary gate. --allow-boundary-fallback is a research override, not part of the measured contract. The quality policy is promoted only for speaker-reference short/medium text; keep using generate_tts.py for centroid and long-form production requests.

This path extracts the windowed speaker embedding once and reuses it for every chunk. Each request gets a new random seed that is reused inside that request; the release gate varies request seeds, so stability does not depend on one global fixed seed. --reference-audio-conditioning tokens enables the raw-reference-token research backend; it is not the production default.

Recommended generation defaults:

  • cfg_value=2.0
  • inference_timesteps=10
  • retry_badcase=False for the measured long-form contract
  • retry_badcase_ratio_threshold=6.0 so retry detection is not a pace cap
  • stop_threshold=0.65, stop_consecutive=2 for offline generation
  • max_len=2000 unless a longer chunk is intentionally needed
  • target-chars-per-sec=4.0 for controlled long-form chunking
  • continuation-context-sec=0 until generated-context continuation passes its own gate
  • Reference wavs used to extract speaker embeddings should be at least 3 seconds.

The hosted interactive demo uses a separate endpoint-only policy with the same weights: a 0.50 stop threshold that relaxes after 75% of the native-rate duration estimate to 0.05 at 95%, one stop hit, and a native-rate hard cap plus one latent step. It applies pace correction after generation. The offline defaults above remain the fixed evaluation contract.

Safety:

  • Use only rights-cleared reference audio or speaker embeddings.
  • Do not present generated speech as a real person or real notification unless that is explicitly authorized and reviewed.