| # Attribution | |
| `s2s-a0-en` contains Mimi-tokenized (kyutai/mimi, CC-BY-4.0) derivatives of: | |
| - **Multilingual LibriSpeech (English)** — via `parler-tts/mls_eng`. CC-BY-4.0. | |
| Pratap et al., "MLS: A Large-Scale Multilingual Dataset for Speech Research" (2020). | |
| - **LibriTTS-R** — via `blabble-io/libritts_r`. CC-BY-4.0. | |
| Koizumi et al., "LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus" (2023). | |
| - **gemini-flash-2.0-speech** — `shb777/gemini-flash-2.0-speech`. Apache-2.0. | |
| Text streams are tokenized with the Qwen3-0.6B tokenizer (Apache-2.0). | |
| This bucket contains no audio waveforms, only discrete codec tokens + token ids. | |
Xet Storage Details
- Size:
- 643 Bytes
- Xet hash:
- 69801b3e29aec919fab19df5e56ba95507258cbda0bbb872410e24c2f0cec804
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.