Tachiwin multilingual tokenizer
A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.
Corpus weighting
| Component | Target |
|---|---|
| Modern + old exotic-language data | 70% |
| English | 10% |
| Spanish | 10% |
| Code | 10% |
The complete available exotic-language corpus is used as the 70% anchor. Its existing modern/old composition is preserved.
Tokenizer
- Model: BPE, trained via
Tokenizer.train_from_iteratorover a streaming line generator (constant memory regardless of corpus size) - Vocabulary target: 64,000
- Initial alphabet: complete ByteLevel alphabet
- ByteLevel GPT-2 regex: disabled
- Unicode normalizer: none
- Special tokens: 282
- Human-language tags: 248
Corpus size
Total materialized corpus: 457,300,912 bytes (0.426 GiB)
Evaluation
The tokenizer was evaluated against the Tachiwin language catalogue. Languages without text samples are skipped. For each language, all available samples are concatenated ONLY within that language for aggregate fertility statistics. Metrics: characters/token, tokens/character, UTF-8 bytes/token, tokens/UTF-8 byte, exact round-trip preservation.
Important training note
The Hugging Face BPE trainer does not expose an internal resumable
merge-state checkpoint. The recipe therefore treats the completed
tokenizer.json as the training checkpoint:
- corpus preparation is resumable (both the exotic streaming pass and the capped external-corpus downloads reuse existing shards);
- recipe/statistics/checksums are stored in
recipe/; - if
tokenizer.jsonalready exists, subsequent runs skip BPE training; - an interrupted BPE computation itself must be restarted.
Repository evaluation artifacts
- evaluation/catalogue.json
- evaluation/language_fertility.csv
- evaluation/language_fertility.json
- evaluation/evaluation_summary.json