Sansar Sanskrit tokenizer
An 8,000-piece SentencePiece tokenizer built for Sanskrit, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models trained from scratch.
- Compact: 2.6 to 3.9 tokens per word with an 8k vocabulary. General-purpose tokenizers with 68k to 200k vocabularies need 2.9 to 8.6 on the same text.
- Never splits a vowel sign from its consonant: 0 orphaned vowel signs per 1,000 tokens (public tokenizers: 94 to 464).
- Lossless: text in, the identical text (NFC) back, on all 9,170 held-out test texts, including the 453 web texts that mix English words or IAST romanisation into Sanskrit.
Quick start
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("MuseMesh/sansar-sanskrit-tokenizer", revision="v0.1.0")
sys.path.insert(0, path)
from sansar_tokenizer import SansarTokenizer
tok = SansarTokenizer()
ids = tok.encode("धर्मक्षेत्रे कुरुक्षेत्रे समवेता युयुत्सवः")
print(len(ids), tok.tokenize("धर्मक्षेत्रे कुरुक्षेत्रे")) # pieces as the model sees them (SLP1)
print(tok.decode(ids)) # the same Devanagari string
Requires sentencepiece. The wrapper converts Devanagari to SLP1 (a lossless ASCII spelling of Devanagari) with the
bundled translit.py, runs the SentencePiece model, and converts back. English words inside the text are fenced
so they come back unchanged. To use the model directly, feed it SLP1:
spm.SentencePieceProcessor(model_file="tokenizer.model").
How it compares
Tokens per whitespace word (lower is better) and orphaned vowel signs per 1,000 tokens, measured on held-out Sanskrit with the same harness for every tokenizer. Sets: classical = DCS gold sentences (3,000), Gītā (700 verses), prose (2,470), Vedic = accented Ṛgveda pādas (1,000), web = out-of-domain web text (2,000).
| tokenizer | vocab | classical | Gītā | prose | Vedic | web | orphans /1k | lossless |
|---|---|---|---|---|---|---|---|---|
| Sansar 8k (this) | 8,000 | 2.82 | 2.79 | 2.58 | 3.85 | 3.19 | 0 to 0.15 | yes |
| Sarvam-1 | 68,096 | 3.57 | 3.66 | 2.94 | 5.36 | 3.56 | 94 to 223 | yes |
| GPT-4o (o200k_base) | 200,019 | 3.91 | 3.62 | 3.18 | 4.93 | 3.73 | 236 to 439 | yes |
| DeepSeek-V3 | 128,815 | 4.98 | 4.60 | 4.12 | 5.71 | 4.91 | 292 to 464 | yes |
| Qwen3 | 151,669 | 7.84 | 7.56 | 6.65 | 7.56 | 7.89 | 252 to 333 | yes |
| GPT-4 (cl100k_base) | 100,277 | 8.52 | 8.30 | 7.22 | 8.25 | 8.56 | 231 to 307 | yes |
| IndicBERTv2 | 250,000 | 2.43 | 2.34 | 2.00 | 2.85 | 2.40 | 198 to 237 | no (every text changes) |
IndicBERTv2 is shorter only because its normalisation rewrites the text; it fails the round trip on all 9,170 texts.
Per-set tables with more metrics (characters per token, Rényi efficiency, tokens per DCS gold word) are in
eval/results.md. The test texts themselves are not published: they are the frozen held-out sets of the Sansar
models.
How it was built
- SentencePiece unigram, vocab 8,000, trained on SLP1. Settings:
max_sentencepiece_length=32,normalization_rule_name=identity(NFC applied beforehand),remove_extra_whitespaces=false(verse layout survives),byte_fallback=true,character_coverage=1.0,user_defined_symbols=।॥'(daṇḍa, double daṇḍa, avagraha),max_sentence_length=1048576(the default 4192 silently drops two thirds of a corpus of long Sanskrit records). Full config:training_config.json. - Why SLP1: in our experiments (unigram/BPE × 4k to 32k × Devanagari/SLP1) the transliteration costs nothing on compression, removes all orphaned vowel signs, and stores text in 2.4× fewer bytes. Unigram beat BPE by 2.8% bits per byte in matched 20M-parameter language models; 8k and 16k tied, 32k lost.
- Training text: 1,437,305 lines, 1.01 GB of Devanagari from the permissive and share-alike tiers of the Sansar corpus (classical e-texts, treebanks, Wikisource and Wikipedia, dictionaries). Every held-out test text (9,655 keys) was removed before training and the exclusion was verified (0 leaks).
- In use: the frozen tokenizer of every Sansar Sanskrit model so far (20M to 125M parameters).
Limitations
- Vedic accents fragment pieces: Vedic is the weakest set (3.85 tokens/word). Accent handling is the next planned change.
- English words and IAST inside Sanskrit text round trip through the wrapper (fenced), but each run costs two extra marks of byte-fallback tokens. Rare Vedic Extended signs also pass through as byte-fallback tokens.
- Built for Sanskrit. Hindi and other Devanagari languages tokenize, but less efficiently.
Versions
| version | date | change |
|---|---|---|
| v0.1.0 | 2026-10-01 | first release; see CHANGELOG.md |
Every release is a git tag: pass revision="v0.1.0" to pin one.
Licence
This release is for research and non-commercial use. A commercially licensed version, trained only on permissively licensed text, is planned as a separate release with its own version line.
tokenizer.model,tokenizer.vocab: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited".- Code (
sansar_tokenizer.py,translit.py,test_roundtrip.py): Apache-2.0 (LICENSE-CODE). - Commercial use: contact kushal@muse-mesh.com.
Citation
@misc{sansar_tokenizer_2026,
title = {Sansar Sanskrit tokenizer},
author = {Muse Mesh},
year = {2026},
note = {SentencePiece unigram 8k over SLP1, v0.1.0},
url = {https://huggingface.co/MuseMesh/sansar-sanskrit-tokenizer}
}
Contact: kushal@muse-mesh.com