LingkodAI Whisper: Cebuano

Speech-to-text for Cebuano, fine-tuned from openai/whisper-medium. It transcribes queries identified as Cebuano by the audio-based language identification stage of the LingkodAI pipeline.

How it works

Whisper doesn't support Cebuano out of the box, so training added a new language token, <|cebuano|>, at id 51865.

Only the decoder was fine-tuned; the encoder was kept frozen, so it is identical to the original whisper-medium encoder.

Training data

Cebuano read speech from the Philippine Languages Database (PLD), and resampled to 16 kHz, with a speaker-independent train/validation/test split.

Usage

The <|cebuano|> token must be forced at the start of decoding. Without it, Whisper will guess a different language.

import torch, librosa
from transformers import WhisperProcessor, WhisperForConditionalGeneration

repo = "lingkodai/lingkodai-whisper-ceb"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32

processor = WhisperProcessor.from_pretrained(repo, task="transcribe")
model = WhisperForConditionalGeneration.from_pretrained(repo, torch_dtype=dtype).to(device).eval()

# Sanity checks: tokenizer and weights must match
assert len(processor.tokenizer) == model.get_input_embeddings().weight.shape[0]
ceb_id = processor.tokenizer.convert_tokens_to_ids("<|cebuano|>")
assert ceb_id == 51865

# Force: <|cebuano|> <|transcribe|> <|notimestamps|>
model.generation_config.forced_decoder_ids = [
    [1, ceb_id],
    [2, processor.tokenizer.convert_tokens_to_ids("<|transcribe|>")],
    [3, processor.tokenizer.convert_tokens_to_ids("<|notimestamps|>")],
]

audio, _ = librosa.load("clip.wav", sr=16000)
features = processor.feature_extractor(audio, sampling_rate=16000, return_tensors="pt").input_features.to(device, dtype=dtype)
with torch.no_grad():
    ids = model.generate(features, max_length=225)
print(processor.tokenizer.batch_decode(ids, skip_special_tokens=True)[0].strip())
Downloads last month
70
Safetensors
Model size
0.8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for beabucayan/lingkodai-whisper-ceb

Finetuned
(960)
this model