edwixx/s2s-a0-en / ATTRIBUTION.md
edwixx's picture
|
download
raw
643 Bytes
# Attribution
`s2s-a0-en` contains Mimi-tokenized (kyutai/mimi, CC-BY-4.0) derivatives of:
- **Multilingual LibriSpeech (English)** — via `parler-tts/mls_eng`. CC-BY-4.0.
Pratap et al., "MLS: A Large-Scale Multilingual Dataset for Speech Research" (2020).
- **LibriTTS-R** — via `blabble-io/libritts_r`. CC-BY-4.0.
Koizumi et al., "LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus" (2023).
- **gemini-flash-2.0-speech**`shb777/gemini-flash-2.0-speech`. Apache-2.0.
Text streams are tokenized with the Qwen3-0.6B tokenizer (Apache-2.0).
This bucket contains no audio waveforms, only discrete codec tokens + token ids.

Xet Storage Details

Size:
643 Bytes
·
Xet hash:
69801b3e29aec919fab19df5e56ba95507258cbda0bbb872410e24c2f0cec804

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.