Semantic-VAD whisper-small

An audio-only end-of-turn detector for voice agents. Give it the last 8 s of the caller's audio at 16 kHz and it returns the probability that the caller has finished speaking. It needs no transcript, so it can answer the moment a pause begins.

It is a Whisper-small encoder (88 M parameters) with a small head, shipped as int8 ONNX (95 MB) for one CPU thread and as PyTorch weights.

Results

The charts below put every model of the family next to the open detectors, at the budgets LiveKit's eot-bench reports: the fewest interruptions a model can reach within 300 ms and 600 ms of average waiting, and the shortest wait it can reach while interrupting at most 5 % and 10 % of pauses.

Best false-cutoff rate at a 300 / 600 ms latency budget, eot-bench-data

Best false-cutoff rate at a 300 / 600 ms latency budget, telephony test

Best mean latency at a 5 / 10 % false-cutoff budget, eot-bench-data

Best mean latency at a 5 / 10 % false-cutoff budget, telephony test

ROC AUC, eot-bench-data

ROC AUC, telephony test

Every operating point

Each curve is the best trade a policy can reach: how long a finished turn waits against how many pauses get interrupted. Lower left is better.

Latency against cut-offs on eot-bench-data

Latency against cut-offs on the telephony test

Per language on eot-bench-data, the new model is better than the 2026-09-07 release in 11 of the 14 languages:

False cut-offs at 300 ms per language

What changed on 2026-10-03

This revision replaces the release of 2026-09-07, which is kept on the branch release-2026-09-07. It is much better on LiveKit's multilingual set and a little worse on our telephony test:

telephony test: cut-offs at 300 ms telephony: latency at 5 % telephony AUC eot-bench-data: cut-offs at 300 ms eot-bench-data: latency at 5 % eot-bench-data AUC
this model, int8 ONNX 43.4 % 1,803 ms 0.850 20.1 % 801 ms 0.915
this model, PyTorch 43.4 % 1,793 ms 0.850 19.2 % 799 ms 0.917
previous release, int8 ONNX 42.1 % 1,750 ms 0.858 23.6 % 846 ms 0.897
LiveKit Turn Detector v1 not run 14.8 % 688 ms 0.941
LiveKit Turn Detector v1-mini, audio 61.6 % 2,043 ms 0.748 29.9 % 924 ms 0.840
SmartTurn v3.2, CPU int8 69.6 % 2,274 ms 0.670 39.0 % 914 ms 0.789
Scicom Semantic VAD, large, GPU 40.7 % 1,779 ms 0.861 15.8 % 718 ms 0.937
silence timer only 77.6 % 2,020 ms 54.9 % 1,100 ms

Lower is better for cut-offs and latency, higher for AUC.

  • Cut-offs at 300 ms: the share of mid-turn pauses the agent would interrupt if it may wait 300 ms on average after a finished turn. Latency at 5 %: the mean wait after a finished turn if at most 5 % of pauses may be interrupted. AUC at 0.2 s into each pause. All three come from LiveKit's eot-bench: every silence of at least 0.1 s, every policy of threshold, action delay and timeout, the best policy at each budget.
  • eot-bench-data: LiveKit's own 14-language set (livekit/eot-bench-data, validation), scored on the same span set as LiveKit's published runs.
  • The previous release's telephony numbers are slightly flattering: its checkpoint was chosen by early stopping on a sample of the telephony test split. This model never saw that split before it was scored.

In short: 3.5 points fewer interruptions on LiveKit's set and 45 ms less waiting there, for 1.3 points more on telephony. If your traffic is telephony only, the previous release (revision="release-2026-09-07") is still slightly better there.

Usage

ONNX on one CPU thread, no torch needed:

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
from transformers import WhisperFeatureExtractor

REPO, SR, WINDOW = "Scicom-intl/semantic-vad-eot-whisper-small", 16000, 8 * 16000
opts = ort.SessionOptions(); opts.intra_op_num_threads = 1
sess = ort.InferenceSession(hf_hub_download(REPO, "onnx/model.int8.onnx"), opts, providers=["CPUExecutionProvider"])
fe = WhisperFeatureExtractor(feature_size=80, sampling_rate=SR, chunk_length=8)

def p_end_of_turn(pcm: np.ndarray) -> float:
    """pcm: float32 in [-1, 1] at 16 kHz, the caller's audio up to now (any length)."""
    pcm = np.asarray(pcm, dtype=np.float32)
    if pcm.size and np.abs(pcm).max() > 1.5:   # int16-scale samples to unit float
        pcm = pcm / 32768.0
    pcm = pcm[-WINDOW:] if len(pcm) >= WINDOW else np.pad(pcm, (WINDOW - len(pcm), 0))
    feats = fe([pcm], sampling_rate=SR, return_tensors="np", padding="max_length", max_length=WINDOW,
               truncation=True, do_normalize=False)["input_features"].astype(np.float32)
    return float(sess.run(None, {"input_features": feats})[0].reshape(-1)[0])

Threshold: 0.5. eot-bench's best policies for this model use 0.54 to 0.58 at the 300 ms budget, close to the previous release's 0.57 to 0.60.

LiveKit Agents: wrap p_end_of_turn as the predict(pcm) backend of STT-API's SemanticVAD turn detector. LiveKit hands the backend int16-scale samples; the snippet rescales them.

To go back to the previous release, pass revision="release-2026-09-07" to hf_hub_download.

Serving cost

About 185 ms per prediction on one CPU thread (int8, measured at export). A call asks the detector about 0.15 times a second, so one thread per call is plenty, but the work is real: run it in a sidecar or with a few threads rather than inside a busy agent process. A 164-core node sustains about 290 predictions a second with 64 single-thread processes. The cost does not depend on how much audio you send; the window is always 8 s.

Files

file what
onnx/model.int8.onnx MatMul-only dynamic int8, 95 MB; probability differs from PyTorch by 0.023 on average, 0.12 at most
onnx/model.fp32.onnx fp32 export, 350 MB; within 1.4e-6 of PyTorch
onnx/export_report.json sizes, parity against PyTorch, latency at export
encoder/ the encoder with LoRA merged, HF format (config.json, model.safetensors, bf16)
eot_head.pt {"state_dict": LayerNorm, Linear(768, 256), GELU, Linear(256, 1), "pooling": "last5"}
eot_window.json, preprocessor_config.json the input contract: 8 s window, 80 mel bins, 16 kHz, no mel normalisation, mean of the last 5 encoder frames

Input: input_features [batch, 80, 800] float32, the Whisper log-mel of the last 8 s of audio, left-padded with zeros when shorter, do_normalize=False. Output: probability [batch, 1], already through the sigmoid.

License

Apache-2.0, like the Whisper encoder it fine-tunes.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Scicom-intl/semantic-vad-eot-whisper-small

Quantized
(257)
this model

Collection including Scicom-intl/semantic-vad-eot-whisper-small