edwixx/s2s-a0-en / ATTRIBUTION.md
edwixx's picture
|
download
raw
643 Bytes

Attribution

s2s-a0-en contains Mimi-tokenized (kyutai/mimi, CC-BY-4.0) derivatives of:

  • Multilingual LibriSpeech (English) — via parler-tts/mls_eng. CC-BY-4.0. Pratap et al., "MLS: A Large-Scale Multilingual Dataset for Speech Research" (2020).
  • LibriTTS-R — via blabble-io/libritts_r. CC-BY-4.0. Koizumi et al., "LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus" (2023).
  • gemini-flash-2.0-speechshb777/gemini-flash-2.0-speech. Apache-2.0.

Text streams are tokenized with the Qwen3-0.6B tokenizer (Apache-2.0). This bucket contains no audio waveforms, only discrete codec tokens + token ids.

Xet Storage Details

Size:
643 Bytes
·
Xet hash:
69801b3e29aec919fab19df5e56ba95507258cbda0bbb872410e24c2f0cec804

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.