Instructions to use beabucayan/lingkodai-whisper-ceb with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use beabucayan/lingkodai-whisper-ceb with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="beabucayan/lingkodai-whisper-ceb")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("beabucayan/lingkodai-whisper-ceb") model = AutoModelForSpeechSeq2Seq.from_pretrained("beabucayan/lingkodai-whisper-ceb", device_map="auto") - Notebooks
- Google Colab
- Kaggle
LingkodAI Whisper: Cebuano
Speech-to-text for Cebuano, fine-tuned from openai/whisper-medium. It transcribes queries identified as Cebuano by the audio-based language identification stage of the LingkodAI pipeline.
How it works
Whisper doesn't support Cebuano out of the box, so training added a new language token, <|cebuano|>, at id 51865.
Only the decoder was fine-tuned; the encoder was kept frozen, so it is identical to the original whisper-medium encoder.
Training data
Cebuano read speech from the Philippine Languages Database (PLD), and resampled to 16 kHz, with a speaker-independent train/validation/test split.
Usage
The <|cebuano|> token must be forced at the start of decoding. Without it, Whisper will guess a different language.
import torch, librosa
from transformers import WhisperProcessor, WhisperForConditionalGeneration
repo = "lingkodai/lingkodai-whisper-ceb"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32
processor = WhisperProcessor.from_pretrained(repo, task="transcribe")
model = WhisperForConditionalGeneration.from_pretrained(repo, torch_dtype=dtype).to(device).eval()
# Sanity checks: tokenizer and weights must match
assert len(processor.tokenizer) == model.get_input_embeddings().weight.shape[0]
ceb_id = processor.tokenizer.convert_tokens_to_ids("<|cebuano|>")
assert ceb_id == 51865
# Force: <|cebuano|> <|transcribe|> <|notimestamps|>
model.generation_config.forced_decoder_ids = [
[1, ceb_id],
[2, processor.tokenizer.convert_tokens_to_ids("<|transcribe|>")],
[3, processor.tokenizer.convert_tokens_to_ids("<|notimestamps|>")],
]
audio, _ = librosa.load("clip.wav", sr=16000)
features = processor.feature_extractor(audio, sampling_rate=16000, return_tensors="pt").input_features.to(device, dtype=dtype)
with torch.no_grad():
ids = model.generate(features, max_length=225)
print(processor.tokenizer.batch_decode(ids, skip_special_tokens=True)[0].strip())
- Downloads last month
- 70
Model tree for beabucayan/lingkodai-whisper-ceb
Base model
openai/whisper-medium