Instructions to use veeranool/Telugu-vaachakam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use veeranool/Telugu-vaachakam with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="veeranool/Telugu-vaachakam")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("veeranool/Telugu-vaachakam") model = AutoModelForMaskedLM.from_pretrained("veeranool/Telugu-vaachakam", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Telugu-vaachakam
A Telugu-only BERT masked-language model trained from scratch for Telugu language understanding and fine-tuning.
Model summary
- Architecture: BERT masked language model
- Parameters: approximately 268M
- Vocabulary: 64,000 Telugu WordPiece tokens
- Maximum sequence length: 512
- Intended use: masked-token prediction, embeddings, NER, classification, and extractive QA fine-tuning
Measured benchmark
On the google/xtreme PAN-X.te Telugu NER test split, using the same fine-tuning protocol for both models:
| Model | Entity F1 | Precision | Recall | Token accuracy |
|---|---|---|---|---|
| Telugu-vaachakam | 69.71% | 66.06% | 73.78% | 92.32% |
| Public Telugu BERT baseline | 57.42% | 54.67% | 60.46% | 88.09% |
The classification head was fine-tuned on the PAN-X.te training split and evaluated on its separate test split. These results are a benchmark comparison, not a zero-shot evaluation. The baseline is described generically here to keep the model card focused on Telugu-vaachakam rather than another project's identity.
MLM diagnostic
An in-domain masked-language-model diagnostic measured loss 2.5262, perplexity 12.51, and masked-token accuracy 54.30%. This is not a leakage-free held-out score because the pretraining run did not reserve a validation split.
No claims are made here about sentiment, topic classification, QA, or overall state-of-the-art performance because those tasks were not evaluated.
Usage
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("veeranool/Telugu-vaachakam")
model = AutoModelForMaskedLM.from_pretrained("veeranool/Telugu-vaachakam")
text = "తెలుగు [MASK] ద్రావిడ భాష."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
Limitations
This is an encoder-only BERT model, not a text-generation or chat model. Performance outside Telugu and on tasks not listed above has not been established. The model may reflect biases or noise in its training data.
- Downloads last month
- 25