- π±π° SinhalaLM v0.0.1
- π Table of contents
- π Overview
- π‘ Why SinhalaLM
- π₯ Authors
- π Model summary
- π― Intended use
- ποΈ Architecture
- π€ Tokenizer
- βοΈ Training procedure
- π₯οΈ Training hardware
- π How to use
- β οΈ Limitations and biases
- π Files in this repo
- π License
- π Citation
- π¬ Contact
- π Table of contents
π±π° SinhalaLM v0.0.1
A GPT-style Sinhala Language Model, trained entirely from scratch
SinhalaLM Β· v0.0.1 Β· by HelaAI
π Table of contents
- Overview
- Why SinhalaLM
- Authors
- Model summary
- Intended use
- Architecture
- Tokenizer
- Training procedure
- Training hardware
- How to use
- Limitations and biases
- Files in this repo
- License
- Citation
- Contact
π Overview
SinhalaLM is a GPT-style, decoder-only causal language model trained entirely from scratch on a large Sinhala text corpus. Rather than adapting an existing multilingual model to Sinhala through continual pretraining, SinhalaLM was designed and sized specifically for the Sinhala language, using a compute-aware architecture built around three modern techniques:
- π Rotary Position Embeddings (RoPE) for relative positional encoding
- β‘ SwiGLU gated feed-forward blocks instead of a plain ReLU MLP
- π€ A custom byte-pair encoding (BPE) tokenizer trained specifically on the Sinhala script (and retained English characters) present in the training corpus
This is a custom architecture β it is not a standard transformers (AutoModel) class. The full model definition ships in this repository as modeling_sinhalalm.py and must be imported directly. Full loading and generation instructions are provided below.
Detailed internal architecture of a single SinhalaLM decoder block β showing the QKV split, per-head RoPE application, the fused Flash-Attention kernel, both residual connections, and the SwiGLU feed-forward path with weight tying between the token embedding and the LM head.
β οΈ Private repository. Only Hugging Face accounts explicitly granted access can pull this repo. Do not share the access token.
π‘ Why SinhalaLM
Sinhala is spoken by over 20 million people, yet it remains comparatively under-served by mainstream multilingual language models, which tend to prioritize languages with abundant digital text and stronger commercial incentives. The dominant strategy for extending LLM capability to low-resource languages like Sinhala has been continual pretraining of a large existing model (e.g. adapting an 8-billion-parameter backbone) β an approach that requires substantial GPU memory and incurs high inference cost.
SinhalaLM explores the alternative path: building a compact model from scratch, sized to match the pretraining data actually available for the language. At ~32.9M parameters β roughly 240Γ smaller than an 8B-parameter continually pretrained alternative β SinhalaLM is:
- π° Cheap to train β fits comfortably on a single consumer/prosumer GPU
- π Cheap and fast to deploy β low memory footprint, low inference latency
- π― Fully fine-tunable β its small size allows full fine-tuning (updating every parameter) on a single GPU, rather than relying on low-rank adapters
π₯ Authors
| Name | Affiliation | Contact |
|---|---|---|
| Chamara Vishwajith | University of Ruhuna, Sri Lanka | rajapaksha_rmcv_e23@engug.ruh.ac.lk, cam.chamara@giam.com |
| Dr. Kushan Sudheera | Dept. of Electrical and Information Engineering, University of Ruhuna, Sri Lanka | kushan@eie.ruh.ac.lk |
| Dineth Jayakody | Old Dominion University, United States | djaya003@odu.edu |
π Model summary
| Model name | SinhalaLM |
| Version | v0.0.1 |
| Developed by | HelaAI |
| Model type | Decoder-only causal language model (GPT-style) |
| Language | Sinhala (si), with retained English letters/punctuation |
| Training approach | From-scratch pretraining (not continually pretrained from an existing checkpoint) |
| License | Other β see License |
π― Intended use
SinhalaLM is intended for Sinhala text generation and continuation β research, prototyping, and experimentation with a from-scratch Sinhala language model. It has also been fine-tuned and evaluated as a base for downstream Sinhala tasks such as text classification and question answering.
Not intended for:
- β Production use without further evaluation, safety review, and human oversight.
- β Factual question-answering out of the box β the base model is a language model, not an instruction-tuned or fact-checked assistant, and will generate fluent but not necessarily accurate or true text.
- β Any use where generated text could cause harm if biased, incoherent, or false (e.g. medical, legal, or financial guidance).
ποΈ Architecture
SinhalaLM is a decoder-only transformer, built with a pre-normalized block structure and several architectural choices intended to improve training stability, throughput, and length generalization relative to a vanilla GPT baseline.
| Component | Design choice |
|---|---|
| Parameters | ~32.9M |
| Embedding dim | 512 |
| Attention heads | 8 |
| Transformer layers | 8 |
| Context length | 512 tokens |
| Vocabulary size | 15,000 |
| Position encoding | Rotary (RoPE) |
| Feed-forward | SwiGLU (gated activation) |
| Normalization | RMSNorm (pre-norm, at both sublayers and final output) |
| Attention implementation | PyTorch scaled_dot_product_attention (Flash Attention where available) |
| Weight tying | Input token embedding tied to the output LM head |
| Dropout | 0.2 |
Per-block structure:
- RMSNorm β causal self-attention (RoPE applied to Q/K per head, fused Flash-Attention kernel) β residual add
- RMSNorm β SwiGLU feed-forward (three-projection gated MLP) β residual add
Systems-level details:
- Attention: Implemented with a fused scaled-dot-product-attention kernel, avoiding materialization of the full
T Γ Tattention matrix and improving training throughput. - Output layer: Weight tying shares parameters between the token embeddings and the language-modeling head, reducing the cost of the vocabulary projection.
- Precision & compute: Trained in
bfloat16mixed precision with gradient accumulation over microbatches; the model graph is compiled ahead of time for operation fusion and improved GPU utilization.
For the full annotated diagram (query/key/value split, per-head RoPE, fused attention kernel, residual paths, and the SwiGLU projections), see the architecture figure above.
π€ Tokenizer
A custom byte-pair encoding (BPE) tokenizer, trained from scratch on the training corpus via a streaming, multi-process word-frequency counting pass (so the corpus is never fully materialized in memory), followed by an iterative greedy merge procedure.
- Preserves Sinhala script (U+0D80βU+0DFF), English letters, digits, and common punctuation β nothing outside that set survives cleaning.
- Special tokens:
<PAD>,<UNK>,<BOS>,<EOS>. - Vocabulary size: 15,000 tokens.
- The authoritative loader is the
FastBPETokenizerclass shipped inmodeling_sinhalalm.py; HF-stylevocab.json/merges.txt/tokenizer.jsonfiles are also included for interoperability and inspection. - The trained tokenizer is used to pre-encode the corpus once into a flat, memory-mapped binary token cache; training and validation batches are drawn from this cache via random-offset sampling using a pinned-memory, prefetching data loader β allowing training to scale to corpora substantially larger than available system RAM.
βοΈ Training procedure
Pretraining corpus: ~304 million Sinhala tokens, derived from the Sinhala Text Dataset.
| Hyperparameter | Value |
|---|---|
| Micro batch size | 32 |
| Gradient accumulation steps | 4 |
| Effective batch size | 128 |
| Max iterations (configured) | 150,000 |
| Warm-up iterations | 3,000 |
| Peak learning rate | 3 Γ 10β»β΄ |
| Min learning rate | 3 Γ 10β»β΅ |
| LR schedule | Cosine decay with linear warmup |
| Weight decay | 0.15 |
| Optimizer | AdamW (Ξ² = (0.9, 0.95)) |
| Precision | torch.bfloat16 |
| Gradient clipping | Global norm 1.0 |
Fine-tuning: Because SinhalaLM has only ~32.9M parameters, it can be fully fine-tuned (every weight updated, no low-rank adapter constraint) on a single GPU. This was used to adapt SinhalaLM for downstream Sinhala tasks using an Alpaca-style supervised prompt format that separates instruction, question, and expected response with textual markers (e.g. for question answering: an instruction line, a ### ΰΆ΄ΰ·βΰΆ»ΰ·ΰ·ΰΆ±ΰΆΊ: question section, and a ### ΰΆ΄ΰ·ΰ·
ΰ·ΰΆΰ·ΰΆ»: answer section, wrapped in <BOS> / <EOS>).
| Fine-tuning hyperparameter | Value |
|---|---|
| Batch size (grad. accum.) | 8 (8) |
| Epochs | 3 |
| Peak learning rate | 1 Γ 10β»β΄ |
π₯οΈ Training hardware
| Device | CUDA |
| GPU | NVIDIA GeForce RTX 5090 |
| VRAM | 33.7 GB |
| Precision | torch.bfloat16 |
| Data loader workers | 30 |
Full training curves (train/validation loss, learning-rate schedule, throughput, perplexity, and a combined dashboard) plus raw CSV logs are included in the analysis/ folder of this repo:
01_train_loss.png,02_val_loss.png,03_train_vs_val_loss.png04_learning_rate.png,05_throughput.png06_val_loss_and_lr.png,07_perplexity.png,08_smoothed_loss.png09_dashboard.pngβ all of the above in one viewtrain_loss.csv,val_loss.csv,lr_schedule.csv,throughput.csvtraining_summary.txt/training_summary.json
π How to use
Install dependencies:
pip install torch huggingface_hub safetensors
Download, load, and generate:
from huggingface_hub import snapshot_download
import torch, json, sys
# 1. Download the repo (private β requires a token with read access)
local_dir = snapshot_download(repo_id="ChamaraVishwajithRajapaksha/SinhalaLM-V.0.0.1", token="YOUR_HF_TOKEN")
# 2. Import the model + tokenizer classes shipped in this repo
sys.path.insert(0, local_dir)
from modeling_sinhalalm import SinhalaGPT, FastBPETokenizer
# 3. Load config
config = json.load(open(f"{local_dir}/config.json"))
# 4. Load tokenizer
tokenizer = FastBPETokenizer.load(f"{local_dir}/bpe_vocabulary.json")
# 5. Build model + load weights
from safetensors.torch import load_file
state_dict = load_file(f"{local_dir}/model.safetensors")
model = SinhalaGPT(
vocab_size=config["vocab_size"],
embed_dim=config["n_embed"],
n_heads=config["n_heads"],
n_layers=config["n_layers"],
block_size=config["block_size"],
dropout=0.0, # no dropout at inference
)
model.load_state_dict(state_dict)
model.eval()
# 6. Generate (unconditional β starts from <BOS>)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
ctx = torch.tensor([[tokenizer.stoi["<BOS>"]]], dtype=torch.long, device=device)
out = model.generate(ctx, max_new_tokens=200, temperature=0.8, top_k=200)
print(tokenizer.decode(out[0].tolist()))
# 6b. Or continue a prompt instead of generating unconditionally
prompt_ids = tokenizer.encode("ΰΆΰΆΊΰ·ΰΆΆΰ·ΰ·ΰΆ±ΰ·")
eos_id = tokenizer.stoi["<EOS>"]
if prompt_ids and prompt_ids[-1] == eos_id:
prompt_ids = prompt_ids[:-1] # drop auto-appended <EOS> before continuing
ctx = torch.tensor([prompt_ids], dtype=torch.long, device=device)
out = model.generate(ctx, max_new_tokens=200, temperature=0.8, top_k=200)
print(tokenizer.decode(out[0].tolist()))
Generation parameters
| Parameter | Effect |
|---|---|
temperature |
Lower (e.g. 0.5β0.7) β more deterministic/repetitive. Higher (e.g. 1.0+) β more random/creative. |
top_k |
Restricts sampling to the k most likely tokens at each step. Lower = safer, less diverse. |
max_new_tokens |
Number of tokens to generate. Model's max context is 512 tokens total (prompt + generation; older tokens are truncated once exceeded). |
β οΈ Limitations and biases
- Trained on a general Sinhala text corpus without curated filtering for toxicity, bias, or factual accuracy β treat all outputs as unverified.
- As a base (non-instruction-tuned) model, the pretrained checkpoint continues text rather than following instructions or answering questions directly; instruction-following behavior requires fine-tuning.
- May produce repetitive or incoherent text for long generations, and can reflect biases present in the training data.
- Context window is limited to 512 tokens β not suitable for long-document tasks in its current form.
- No safety filtering or content moderation has been applied to outputs.
- SinhalaLM is monolingual by design and does not support multilingual or cross-lingual use cases.
π Files in this repo
| File | Description |
|---|---|
model.safetensors |
Model weights, in the safetensors format |
config.json |
Architecture config + model metadata |
bpe_vocabulary.json |
Full custom tokenizer (vocab + merges + merge priorities) β load with FastBPETokenizer.load(...) |
vocab.json, merges.txt, tokenizer.json, tokenizer_config.json, special_tokens_map.json |
HF-style tokenizer files, for interoperability/inspection (the authoritative loader is FastBPETokenizer) |
modeling_sinhalalm.py |
Full source of the SinhalaGPT model class and FastBPETokenizer |
images/sinhalalm_detailed_architecture.svg |
Detailed architecture diagram of a single SinhalaLM decoder block |
analysis/ |
Training graphs, CSV logs, and text/JSON training summary |
generated_sinhala_text.txt |
Sample output generated at the end of training, if available |
π License
This model is released under a custom/other license. Update this section with your actual terms (e.g. permitted uses, redistribution, commercial use) before granting access to others.
π Citation
If you use this model, please cite:
@misc{sinhalalm001,
title = {SinhalaLM v0.0.1},
author = {Chamara Vishwajith and Dr. Kushan Sudheera and Dineth Jayakody},
year = {2026},
note = {Custom GPT-style Sinhala language model, trained from scratch with RoPE and SwiGLU}
}
π¬ Contact
Developed by HelaAI. For questions or access requests, contact one of the authors listed above, or the repository owner on Hugging Face.
ββ SinhalaLM v0.0.1 Β· HelaAI ββ
- Downloads last month
- 181