Instructions to use Scicom-intl/semantic-vad-eot-whisper-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Scicom-intl/semantic-vad-eot-whisper-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Scicom-intl/semantic-vad-eot-whisper-base")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Scicom-intl/semantic-vad-eot-whisper-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Scicom-intl/semantic-vad-eot-whisper-base", device_map="auto")Semantic-VAD whisper-base
An audio-only end-of-turn detector for voice agents. Give it the last 8 s of the caller's audio at 16 kHz and it returns the probability that the caller has finished speaking. It needs no transcript, so it can answer the moment a pause begins.
It is a Whisper-base encoder (20 M parameters) with a small head, shipped as int8 ONNX (24 MB) for one CPU thread and as PyTorch weights.
Results
The charts below put every model of the family next to the open detectors, at the budgets LiveKit's eot-bench reports: the fewest interruptions a model can reach within 300 ms and 600 ms of average waiting, and the shortest wait it can reach while interrupting at most 5 % and 10 % of pauses.
Every operating point
Each curve is the best trade a policy can reach: how long a finished turn waits against how many pauses get interrupted. Lower left is better.
Per language on eot-bench-data, the new model's int8 export is better than the 2026-09-07 release's in 9 of the 14 languages:
What changed on 2026-10-03
This revision replaces the release of 2026-09-07, which is kept on the branch release-2026-09-07. In PyTorch it is better on both test sets. As
shipped, in int8, it is better on LiveKit's multilingual set and a point worse on our telephony test:
| telephony test: cut-offs at 300 ms | telephony: latency at 5 % | telephony AUC | eot-bench-data: cut-offs at 300 ms | eot-bench-data: latency at 5 % | eot-bench-data AUC | |
|---|---|---|---|---|---|---|
| this model, int8 ONNX | 44.2 % | 1,803 ms | 0.847 | 26.5 % | 861 ms | 0.879 |
| this model, PyTorch | 43.9 % | 1,785 ms | 0.848 | 25.0 % | 858 ms | 0.885 |
| previous release, int8 ONNX | 43.2 % | 1,803 ms | 0.850 | 28.3 % | 874 ms | 0.870 |
| previous release, PyTorch | 44.0 % | 1,810 ms | 0.847 | 27.2 % | 858 ms | 0.877 |
| LiveKit Turn Detector v1 | not run | 14.8 % | 688 ms | 0.941 | ||
| LiveKit Turn Detector v1-mini, audio | 61.6 % | 2,043 ms | 0.748 | 29.9 % | 924 ms | 0.840 |
| SmartTurn v3.2, CPU int8 | 69.6 % | 2,274 ms | 0.670 | 39.0 % | 914 ms | 0.789 |
| Scicom Semantic VAD, large, GPU | 40.7 % | 1,779 ms | 0.861 | 15.8 % | 718 ms | 0.937 |
| silence timer only | 77.6 % | 2,020 ms | 54.9 % | 1,100 ms |
Lower is better for cut-offs and latency, higher for AUC.
- Cut-offs at 300 ms: the share of mid-turn pauses the agent would interrupt if it may wait 300 ms on average after a finished turn. Latency at 5 %: the mean wait after a finished turn if at most 5 % of pauses may be interrupted. AUC at 0.2 s into each pause. All three come from LiveKit's eot-bench: every silence of at least 0.1 s, every policy of threshold, action delay and timeout, the best policy at each budget.
- eot-bench-data: LiveKit's own 14-language set (
livekit/eot-bench-data, validation), scored on the same span set as LiveKit's published runs. - The previous release's telephony numbers are slightly flattering: its checkpoint was chosen by early stopping on a sample of the telephony test split. This model never saw that split before it was scored.
In short, as shipped in int8: 1.8 points fewer interruptions on LiveKit's set and 13 ms less waiting there, for 1.0
point more on telephony. In PyTorch the new model is ahead on both sets; the previous release's int8 export happened to
score better on telephony than its own PyTorch weights. If your traffic is telephony only, the previous release
(revision="release-2026-09-07") is still slightly better there.
Usage
ONNX on one CPU thread, no torch needed:
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
from transformers import WhisperFeatureExtractor
REPO, SR, WINDOW = "Scicom-intl/semantic-vad-eot-whisper-base", 16000, 8 * 16000
opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
sess = ort.InferenceSession(hf_hub_download(REPO, "onnx/model.int8.onnx"), opts, providers=["CPUExecutionProvider"])
fe = WhisperFeatureExtractor(feature_size=80, sampling_rate=SR, chunk_length=8)
def p_end_of_turn(pcm: np.ndarray) -> float:
"""pcm: float32 in [-1, 1] at 16 kHz, the caller's audio up to now (any length)."""
pcm = np.asarray(pcm, dtype=np.float32)
if pcm.size and np.abs(pcm).max() > 1.5: # int16-scale samples to unit float
pcm = pcm / 32768.0
pcm = pcm[-WINDOW:] if len(pcm) >= WINDOW else np.pad(pcm, (WINDOW - len(pcm), 0))
feats = fe([pcm], sampling_rate=SR, return_tensors="np", padding="max_length", max_length=WINDOW,
truncation=True, do_normalize=False)["input_features"].astype(np.float32)
return float(sess.run(None, {"input_features": feats})[0].reshape(-1)[0])
Threshold: 0.5. eot-bench's best policies for this model's int8 export use 0.50 to 0.56 at the 300 ms budget, close to the previous release's 0.49 to 0.53.
LiveKit Agents: wrap p_end_of_turn as the predict(pcm) backend of
STT-API's SemanticVAD turn detector. LiveKit hands the
backend int16-scale samples; the snippet rescales them.
To go back to the previous release, pass revision="release-2026-09-07" to hf_hub_download.
Serving cost
About 56 ms per prediction on one CPU thread (int8, measured at export). A call asks the detector about 0.15 times a second, so about 1 % of one core per call. A 164-core node sustains about 900 predictions a second with 64 single-thread processes, and a 4-core node about 64. The cost does not depend on how much audio you send; the window is always 8 s.
Files
| file | what |
|---|---|
onnx/model.int8.onnx |
MatMul-only dynamic int8, 24 MB; probability differs from PyTorch by 0.023 on average, 0.067 at most |
onnx/model.fp32.onnx |
fp32 export, 81 MB; within 7e-7 of PyTorch |
onnx/export_report.json |
sizes, parity against PyTorch, latency at export |
encoder/ |
the fine-tuned encoder, HF format (config.json, model.safetensors) |
eot_head.pt |
{"state_dict": LayerNorm, Linear(512, 256), GELU, Linear(256, 1), "pooling": "last5"} |
eot_window.json, preprocessor_config.json |
the input contract: 8 s window, 80 mel bins, 16 kHz, no mel normalisation, mean of the last 5 encoder frames |
Input: input_features [batch, 80, 800] float32, the Whisper log-mel of the last 8 s of audio, left-padded with zeros
when shorter, do_normalize=False. Output: probability [batch, 1], already through the sigmoid.
License
Apache-2.0, like the Whisper encoder it fine-tunes.
Model tree for Scicom-intl/semantic-vad-eot-whisper-base
Base model
openai/whisper-base








# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("audio-classification", model="Scicom-intl/semantic-vad-eot-whisper-base")