MuseMesh/sansar-700m
Text Generation • 0.7B • Updated • 54 • 1
Sanskrit-only models trained from scratch, their tokenizer and the open Sanskrit corpus.
Note Best model: ex-Gita 0.5547 bits per byte
Note 252M words, one config per licence
Note 8k SLP1 unigram + Devanagari wrapper
Note Printed Sanskrit page OCR (Qwen3.5-2B fine-tune); 1.05% median letter error vs Google Vision 1.69%. Trained only on clean e-texts, born-digital text layers and synthetic pages.