Sansar Sanskrit tokenizer

An 8,000-piece SentencePiece tokenizer built for Sanskrit, from the Sansar project at Muse Mesh (Hugging Face): Sanskrit language models trained from scratch.

  • Compact: 2.6 to 3.9 tokens per word with an 8k vocabulary. General-purpose tokenizers with 68k to 200k vocabularies need 2.9 to 8.6 on the same text.
  • Never splits a vowel sign from its consonant: 0 orphaned vowel signs per 1,000 tokens (public tokenizers: 94 to 464).
  • Lossless: text in, the identical text (NFC) back, on all 9,170 held-out test texts, including the 453 web texts that mix English words or IAST romanisation into Sanskrit.

Quick start

from huggingface_hub import snapshot_download
import sys
path = snapshot_download("MuseMesh/sansar-sanskrit-tokenizer", revision="v0.1.0")
sys.path.insert(0, path)
from sansar_tokenizer import SansarTokenizer

tok = SansarTokenizer()
ids = tok.encode("धर्मक्षेत्रे कुरुक्षेत्रे समवेता युयुत्सवः")
print(len(ids), tok.tokenize("धर्मक्षेत्रे कुरुक्षेत्रे"))   # pieces as the model sees them (SLP1)
print(tok.decode(ids))          # the same Devanagari string

Requires sentencepiece. The wrapper converts Devanagari to SLP1 (a lossless ASCII spelling of Devanagari) with the bundled translit.py, runs the SentencePiece model, and converts back. English words inside the text are fenced so they come back unchanged. To use the model directly, feed it SLP1: spm.SentencePieceProcessor(model_file="tokenizer.model").

How it compares

Tokens per whitespace word (lower is better) and orphaned vowel signs per 1,000 tokens, measured on held-out Sanskrit with the same harness for every tokenizer. Sets: classical = DCS gold sentences (3,000), Gītā (700 verses), prose (2,470), Vedic = accented Ṛgveda pādas (1,000), web = out-of-domain web text (2,000).

tokenizer vocab classical Gītā prose Vedic web orphans /1k lossless
Sansar 8k (this) 8,000 2.82 2.79 2.58 3.85 3.19 0 to 0.15 yes
Sarvam-1 68,096 3.57 3.66 2.94 5.36 3.56 94 to 223 yes
GPT-4o (o200k_base) 200,019 3.91 3.62 3.18 4.93 3.73 236 to 439 yes
DeepSeek-V3 128,815 4.98 4.60 4.12 5.71 4.91 292 to 464 yes
Qwen3 151,669 7.84 7.56 6.65 7.56 7.89 252 to 333 yes
GPT-4 (cl100k_base) 100,277 8.52 8.30 7.22 8.25 8.56 231 to 307 yes
IndicBERTv2 250,000 2.43 2.34 2.00 2.85 2.40 198 to 237 no (every text changes)

IndicBERTv2 is shorter only because its normalisation rewrites the text; it fails the round trip on all 9,170 texts. Per-set tables with more metrics (characters per token, Rényi efficiency, tokens per DCS gold word) are in eval/results.md. The test texts themselves are not published: they are the frozen held-out sets of the Sansar models.

How it was built

  • SentencePiece unigram, vocab 8,000, trained on SLP1. Settings: max_sentencepiece_length=32, normalization_rule_name=identity (NFC applied beforehand), remove_extra_whitespaces=false (verse layout survives), byte_fallback=true, character_coverage=1.0, user_defined_symbols = । ॥ ' (daṇḍa, double daṇḍa, avagraha), max_sentence_length=1048576 (the default 4192 silently drops two thirds of a corpus of long Sanskrit records). Full config: training_config.json.
  • Why SLP1: in our experiments (unigram/BPE × 4k to 32k × Devanagari/SLP1) the transliteration costs nothing on compression, removes all orphaned vowel signs, and stores text in 2.4× fewer bytes. Unigram beat BPE by 2.8% bits per byte in matched 20M-parameter language models; 8k and 16k tied, 32k lost.
  • Training text: 1,437,305 lines, 1.01 GB of Devanagari from the permissive and share-alike tiers of the Sansar corpus (classical e-texts, treebanks, Wikisource and Wikipedia, dictionaries). Every held-out test text (9,655 keys) was removed before training and the exclusion was verified (0 leaks).
  • In use: the frozen tokenizer of every Sansar Sanskrit model so far (20M to 125M parameters).

Limitations

  • Vedic accents fragment pieces: Vedic is the weakest set (3.85 tokens/word). Accent handling is the next planned change.
  • English words and IAST inside Sanskrit text round trip through the wrapper (fenced), but each run costs two extra marks of byte-fallback tokens. Rare Vedic Extended signs also pass through as byte-fallback tokens.
  • Built for Sanskrit. Hindi and other Devanagari languages tokenize, but less efficiently.

Versions

version date change
v0.1.0 2026-10-01 first release; see CHANGELOG.md

Every release is a git tag: pass revision="v0.1.0" to pin one.

Licence

This release is for research and non-commercial use. A commercially licensed version, trained only on permissively licensed text, is planned as a separate release with its own version line.

  • tokenizer.model, tokenizer.vocab: CC BY-NC 4.0 (LICENSE). Non-commercial use (research, teaching, non-profit work) with attribution to "Sansar, Muse Mesh Private Limited".
  • Code (sansar_tokenizer.py, translit.py, test_roundtrip.py): Apache-2.0 (LICENSE-CODE).
  • Commercial use: contact kushal@muse-mesh.com.

Citation

@misc{sansar_tokenizer_2026,
  title  = {Sansar Sanskrit tokenizer},
  author = {Muse Mesh},
  year   = {2026},
  note   = {SentencePiece unigram 8k over SLP1, v0.1.0},
  url    = {https://huggingface.co/MuseMesh/sansar-sanskrit-tokenizer}
}

Contact: kushal@muse-mesh.com

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including MuseMesh/sansar-sanskrit-tokenizer