How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("audio-classification", model="Scicom-intl/semantic-vad-eot-whisper-base")
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("Scicom-intl/semantic-vad-eot-whisper-base", device_map="auto")
Quick Links

Semantic-VAD whisper-base

An audio-only end-of-turn detector for voice agents. Give it the last 8 s of the caller's audio at 16 kHz and it returns the probability that the caller has finished speaking. It needs no transcript, so it can answer the moment a pause begins.

It is a Whisper-base encoder (20 M parameters) with a small head, shipped as int8 ONNX (24 MB) for one CPU thread and as PyTorch weights.

Results

The charts below put every model of the family next to the open detectors, at the budgets LiveKit's eot-bench reports: the fewest interruptions a model can reach within 300 ms and 600 ms of average waiting, and the shortest wait it can reach while interrupting at most 5 % and 10 % of pauses.

Best false-cutoff rate at a 300 / 600 ms latency budget, eot-bench-data

Best false-cutoff rate at a 300 / 600 ms latency budget, telephony test

Best mean latency at a 5 / 10 % false-cutoff budget, eot-bench-data

Best mean latency at a 5 / 10 % false-cutoff budget, telephony test

ROC AUC, eot-bench-data

ROC AUC, telephony test

Every operating point

Each curve is the best trade a policy can reach: how long a finished turn waits against how many pauses get interrupted. Lower left is better.

Latency against cut-offs on eot-bench-data

Latency against cut-offs on the telephony test

Per language on eot-bench-data, the new model's int8 export is better than the 2026-09-07 release's in 9 of the 14 languages:

False cut-offs at 300 ms per language

What changed on 2026-10-03

This revision replaces the release of 2026-09-07, which is kept on the branch release-2026-09-07. In PyTorch it is better on both test sets. As shipped, in int8, it is better on LiveKit's multilingual set and a point worse on our telephony test:

telephony test: cut-offs at 300 ms telephony: latency at 5 % telephony AUC eot-bench-data: cut-offs at 300 ms eot-bench-data: latency at 5 % eot-bench-data AUC
this model, int8 ONNX 44.2 % 1,803 ms 0.847 26.5 % 861 ms 0.879
this model, PyTorch 43.9 % 1,785 ms 0.848 25.0 % 858 ms 0.885
previous release, int8 ONNX 43.2 % 1,803 ms 0.850 28.3 % 874 ms 0.870
previous release, PyTorch 44.0 % 1,810 ms 0.847 27.2 % 858 ms 0.877
LiveKit Turn Detector v1 not run 14.8 % 688 ms 0.941
LiveKit Turn Detector v1-mini, audio 61.6 % 2,043 ms 0.748 29.9 % 924 ms 0.840
SmartTurn v3.2, CPU int8 69.6 % 2,274 ms 0.670 39.0 % 914 ms 0.789
Scicom Semantic VAD, large, GPU 40.7 % 1,779 ms 0.861 15.8 % 718 ms 0.937
silence timer only 77.6 % 2,020 ms 54.9 % 1,100 ms

Lower is better for cut-offs and latency, higher for AUC.

  • Cut-offs at 300 ms: the share of mid-turn pauses the agent would interrupt if it may wait 300 ms on average after a finished turn. Latency at 5 %: the mean wait after a finished turn if at most 5 % of pauses may be interrupted. AUC at 0.2 s into each pause. All three come from LiveKit's eot-bench: every silence of at least 0.1 s, every policy of threshold, action delay and timeout, the best policy at each budget.
  • eot-bench-data: LiveKit's own 14-language set (livekit/eot-bench-data, validation), scored on the same span set as LiveKit's published runs.
  • The previous release's telephony numbers are slightly flattering: its checkpoint was chosen by early stopping on a sample of the telephony test split. This model never saw that split before it was scored.

In short, as shipped in int8: 1.8 points fewer interruptions on LiveKit's set and 13 ms less waiting there, for 1.0 point more on telephony. In PyTorch the new model is ahead on both sets; the previous release's int8 export happened to score better on telephony than its own PyTorch weights. If your traffic is telephony only, the previous release (revision="release-2026-09-07") is still slightly better there.

Usage

ONNX on one CPU thread, no torch needed:

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
from transformers import WhisperFeatureExtractor

REPO, SR, WINDOW = "Scicom-intl/semantic-vad-eot-whisper-base", 16000, 8 * 16000
opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
sess = ort.InferenceSession(hf_hub_download(REPO, "onnx/model.int8.onnx"), opts, providers=["CPUExecutionProvider"])
fe = WhisperFeatureExtractor(feature_size=80, sampling_rate=SR, chunk_length=8)

def p_end_of_turn(pcm: np.ndarray) -> float:
    """pcm: float32 in [-1, 1] at 16 kHz, the caller's audio up to now (any length)."""
    pcm = np.asarray(pcm, dtype=np.float32)
    if pcm.size and np.abs(pcm).max() > 1.5:   # int16-scale samples to unit float
        pcm = pcm / 32768.0
    pcm = pcm[-WINDOW:] if len(pcm) >= WINDOW else np.pad(pcm, (WINDOW - len(pcm), 0))
    feats = fe([pcm], sampling_rate=SR, return_tensors="np", padding="max_length", max_length=WINDOW,
               truncation=True, do_normalize=False)["input_features"].astype(np.float32)
    return float(sess.run(None, {"input_features": feats})[0].reshape(-1)[0])

Threshold: 0.5. eot-bench's best policies for this model's int8 export use 0.50 to 0.56 at the 300 ms budget, close to the previous release's 0.49 to 0.53.

LiveKit Agents: wrap p_end_of_turn as the predict(pcm) backend of STT-API's SemanticVAD turn detector. LiveKit hands the backend int16-scale samples; the snippet rescales them.

To go back to the previous release, pass revision="release-2026-09-07" to hf_hub_download.

Serving cost

About 56 ms per prediction on one CPU thread (int8, measured at export). A call asks the detector about 0.15 times a second, so about 1 % of one core per call. A 164-core node sustains about 900 predictions a second with 64 single-thread processes, and a 4-core node about 64. The cost does not depend on how much audio you send; the window is always 8 s.

Files

file what
onnx/model.int8.onnx MatMul-only dynamic int8, 24 MB; probability differs from PyTorch by 0.023 on average, 0.067 at most
onnx/model.fp32.onnx fp32 export, 81 MB; within 7e-7 of PyTorch
onnx/export_report.json sizes, parity against PyTorch, latency at export
encoder/ the fine-tuned encoder, HF format (config.json, model.safetensors)
eot_head.pt {"state_dict": LayerNorm, Linear(512, 256), GELU, Linear(256, 1), "pooling": "last5"}
eot_window.json, preprocessor_config.json the input contract: 8 s window, 80 mel bins, 16 kHz, no mel normalisation, mean of the last 5 encoder frames

Input: input_features [batch, 80, 800] float32, the Whisper log-mel of the last 8 s of audio, left-padded with zeros when shorter, do_normalize=False. Output: probability [batch, 1], already through the sigmoid.

License

Apache-2.0, like the Whisper encoder it fine-tunes.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Scicom-intl/semantic-vad-eot-whisper-base

Quantized
(245)
this model

Collection including Scicom-intl/semantic-vad-eot-whisper-base