fastText Continual Leaf Surgery: 11-Language Multilingual Rehearsal (Table 3)

This model represents our Continual Learning (Leaf Surgery) architecture trained with the 11-language balanced rehearsal buffer (Table 3 in the research paper):

"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts"
University of Moratuwa, Department of Computer Science & Engineering

1. Architecture & Innovation

  • Hierarchical Softmax Leaf Surgery: The Sinhala (si) leaf node in the official Facebook lid.176.bin Huffman tree was split into Sinhala and Pali branches.
  • Multilingual Rehearsal: Fine-tuned on the balanced 11-language Aya dataset buffer (3 target languages in Sinhala script + 8 global/anchor languages: English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit Devanagari) to preserve cross-lingual representations.

2. Benchmark Performance (11-Language Evaluation)

  • WiLI-2018 (Macro-F1): 0.9811 (Accuracy: 98.00%)
  • CommonLID (Macro-F1): 0.9486 (Accuracy: 96.82%)
  • FLORES+ (Macro-F1): 0.9395 (Accuracy: 92.71%)

3. How to Load and Run Inference in Python

from data_pipeline.fasttext_continual.model import ContinualLID

# Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")

# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction, score = model.predict(text)
print(f"Language: {prediction}, Score: {score:.4f}")
# Output: Language: pi, Score: 0.99...
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for script-langid/fasttext-leaf-surgery-11lang

Finetuned
(6)
this model