Instructions to use Scicom-intl/semantic-vad-eot-whisper-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Scicom-intl/semantic-vad-eot-whisper-small with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Scicom-intl/semantic-vad-eot-whisper-small")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Scicom-intl/semantic-vad-eot-whisper-small", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Semantic-VAD whisper-small
An audio-only end-of-turn detector for voice agents. Give it the last 8 s of the caller's audio at 16 kHz and it returns the probability that the caller has finished speaking. It needs no transcript, so it can answer the moment a pause begins.
It is a Whisper-small encoder (88 M parameters) with a small head, shipped as int8 ONNX (95 MB) for one CPU thread and as PyTorch weights.
Results
The charts below put every model of the family next to the open detectors, at the budgets LiveKit's eot-bench reports: the fewest interruptions a model can reach within 300 ms and 600 ms of average waiting, and the shortest wait it can reach while interrupting at most 5 % and 10 % of pauses.
Every operating point
Each curve is the best trade a policy can reach: how long a finished turn waits against how many pauses get interrupted. Lower left is better.
Per language on eot-bench-data, the new model is better than the 2026-09-07 release in 11 of the 14 languages:
What changed on 2026-10-03
This revision replaces the release of 2026-09-07, which is kept on the branch release-2026-09-07. It is much better on LiveKit's multilingual set
and a little worse on our telephony test:
| telephony test: cut-offs at 300 ms | telephony: latency at 5 % | telephony AUC | eot-bench-data: cut-offs at 300 ms | eot-bench-data: latency at 5 % | eot-bench-data AUC | |
|---|---|---|---|---|---|---|
| this model, int8 ONNX | 43.4 % | 1,803 ms | 0.850 | 20.1 % | 801 ms | 0.915 |
| this model, PyTorch | 43.4 % | 1,793 ms | 0.850 | 19.2 % | 799 ms | 0.917 |
| previous release, int8 ONNX | 42.1 % | 1,750 ms | 0.858 | 23.6 % | 846 ms | 0.897 |
| LiveKit Turn Detector v1 | not run | 14.8 % | 688 ms | 0.941 | ||
| LiveKit Turn Detector v1-mini, audio | 61.6 % | 2,043 ms | 0.748 | 29.9 % | 924 ms | 0.840 |
| SmartTurn v3.2, CPU int8 | 69.6 % | 2,274 ms | 0.670 | 39.0 % | 914 ms | 0.789 |
| Scicom Semantic VAD, large, GPU | 40.7 % | 1,779 ms | 0.861 | 15.8 % | 718 ms | 0.937 |
| silence timer only | 77.6 % | 2,020 ms | 54.9 % | 1,100 ms |
Lower is better for cut-offs and latency, higher for AUC.
- Cut-offs at 300 ms: the share of mid-turn pauses the agent would interrupt if it may wait 300 ms on average after a finished turn. Latency at 5 %: the mean wait after a finished turn if at most 5 % of pauses may be interrupted. AUC at 0.2 s into each pause. All three come from LiveKit's eot-bench: every silence of at least 0.1 s, every policy of threshold, action delay and timeout, the best policy at each budget.
- eot-bench-data: LiveKit's own 14-language set (
livekit/eot-bench-data, validation), scored on the same span set as LiveKit's published runs. - The previous release's telephony numbers are slightly flattering: its checkpoint was chosen by early stopping on a sample of the telephony test split. This model never saw that split before it was scored.
In short: 3.5 points fewer interruptions on LiveKit's set and 45 ms less waiting there, for 1.3 points more on
telephony. If your traffic is telephony only, the previous release (revision="release-2026-09-07") is still
slightly better there.
Usage
ONNX on one CPU thread, no torch needed:
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
from transformers import WhisperFeatureExtractor
REPO, SR, WINDOW = "Scicom-intl/semantic-vad-eot-whisper-small", 16000, 8 * 16000
opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
sess = ort.InferenceSession(hf_hub_download(REPO, "onnx/model.int8.onnx"), opts, providers=["CPUExecutionProvider"])
fe = WhisperFeatureExtractor(feature_size=80, sampling_rate=SR, chunk_length=8)
def p_end_of_turn(pcm: np.ndarray) -> float:
"""pcm: float32 in [-1, 1] at 16 kHz, the caller's audio up to now (any length)."""
pcm = np.asarray(pcm, dtype=np.float32)
if pcm.size and np.abs(pcm).max() > 1.5: # int16-scale samples to unit float
pcm = pcm / 32768.0
pcm = pcm[-WINDOW:] if len(pcm) >= WINDOW else np.pad(pcm, (WINDOW - len(pcm), 0))
feats = fe([pcm], sampling_rate=SR, return_tensors="np", padding="max_length", max_length=WINDOW,
truncation=True, do_normalize=False)["input_features"].astype(np.float32)
return float(sess.run(None, {"input_features": feats})[0].reshape(-1)[0])
Threshold: 0.5. eot-bench's best policies for this model use 0.54 to 0.58 at the 300 ms budget, close to the previous release's 0.57 to 0.60.
LiveKit Agents: wrap p_end_of_turn as the predict(pcm) backend of
STT-API's SemanticVAD turn detector. LiveKit hands the
backend int16-scale samples; the snippet rescales them.
To go back to the previous release, pass revision="release-2026-09-07" to hf_hub_download.
Serving cost
About 185 ms per prediction on one CPU thread (int8, measured at export). A call asks the detector about 0.15 times a second, so one thread per call is plenty, but the work is real: run it in a sidecar or with a few threads rather than inside a busy agent process. A 164-core node sustains about 290 predictions a second with 64 single-thread processes. The cost does not depend on how much audio you send; the window is always 8 s.
Files
| file | what |
|---|---|
onnx/model.int8.onnx |
MatMul-only dynamic int8, 95 MB; probability differs from PyTorch by 0.023 on average, 0.12 at most |
onnx/model.fp32.onnx |
fp32 export, 350 MB; within 1.4e-6 of PyTorch |
onnx/export_report.json |
sizes, parity against PyTorch, latency at export |
encoder/ |
the encoder with LoRA merged, HF format (config.json, model.safetensors, bf16) |
eot_head.pt |
{"state_dict": LayerNorm, Linear(768, 256), GELU, Linear(256, 1), "pooling": "last5"} |
eot_window.json, preprocessor_config.json |
the input contract: 8 s window, 80 mel bins, 16 kHz, no mel normalisation, mean of the last 5 encoder frames |
Input: input_features [batch, 80, 800] float32, the Whisper log-mel of the last 8 s of audio, left-padded with zeros
when shorter, do_normalize=False. Output: probability [batch, 1], already through the sigmoid.
License
Apache-2.0, like the Whisper encoder it fine-tunes.
Model tree for Scicom-intl/semantic-vad-eot-whisper-small
Base model
openai/whisper-small







