Telugu-vaachakam

A Telugu-only BERT masked-language model trained from scratch for Telugu language understanding and fine-tuning.

Model summary

  • Architecture: BERT masked language model
  • Parameters: approximately 268M
  • Vocabulary: 64,000 Telugu WordPiece tokens
  • Maximum sequence length: 512
  • Intended use: masked-token prediction, embeddings, NER, classification, and extractive QA fine-tuning

Measured benchmark

On the google/xtreme PAN-X.te Telugu NER test split, using the same fine-tuning protocol for both models:

Model Entity F1 Precision Recall Token accuracy
Telugu-vaachakam 69.71% 66.06% 73.78% 92.32%
Public Telugu BERT baseline 57.42% 54.67% 60.46% 88.09%

The classification head was fine-tuned on the PAN-X.te training split and evaluated on its separate test split. These results are a benchmark comparison, not a zero-shot evaluation. The baseline is described generically here to keep the model card focused on Telugu-vaachakam rather than another project's identity.

MLM diagnostic

An in-domain masked-language-model diagnostic measured loss 2.5262, perplexity 12.51, and masked-token accuracy 54.30%. This is not a leakage-free held-out score because the pretraining run did not reserve a validation split.

No claims are made here about sentiment, topic classification, QA, or overall state-of-the-art performance because those tasks were not evaluated.

Usage

from transformers import AutoTokenizer, AutoModelForMaskedLM

tokenizer = AutoTokenizer.from_pretrained("veeranool/Telugu-vaachakam")
model = AutoModelForMaskedLM.from_pretrained("veeranool/Telugu-vaachakam")

text = "తెలుగు [MASK] ద్రావిడ భాష."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)

Limitations

This is an encoder-only BERT model, not a text-generation or chat model. Performance outside Telugu and on tasks not listed above has not been established. The model may reflect biases or noise in its training data.

Downloads last month
25
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support