SpeechBrain ECAPA VoxCeleb Core ML

This is a reproducible compiled Core ML export of speechbrain/spkrec-ecapa-voxceleb. It produces a local voice embedding for cosine-similarity speaker comparison.

Model

Property Value
Parameters 20.77 million
Format Compiled Core ML, FLOAT16 compute
Compiled size 39.8 MiB
Input SpeechBrain log-mel [1, frames, 80]
Sample rate 16 kHz
Output 192-dimensional L2-normalized embedding
Mel frames 10 to 3,001 (about 0.1 to 30 seconds)
Minimum deployment macOS 15 / iOS 18

The graph includes sentence mean normalization. Audio-to-mel processing stays outside the graph so apps can share one tested streaming frontend between Core ML and MLX.

Files

File Description
SpeechBrainECAPAVoxCeleb.mlmodelc/ Precompiled Core ML model
frontend.py Reproducible SpeechBrain log-mel frontend
config.json Graph, audio frontend, source revision, and checksums
artifact_manifest.json SHA-256 and size of every compiled artifact file
validation.json Conversion parity and local latency measurements
requirements.txt Pinned standalone runtime dependencies
LICENSE Apache 2.0 license

Validation

Check Result Meaning
Minimum PyTorch/Core ML output cosine 0.999995 1.0 is identical direction
Maximum absolute error 0.000702 Lower is better
Warm local inference 5.0 ms Median after warm-up on the conversion machine

The official source reports 0.80% equal-error rate on the cleaned VoxCeleb1 test set. That upstream result is not presented as a new benchmark for this conversion. Before changing a speaker-verification threshold, validate the compiled artifact on VoxCeleb1-O and the microphones and languages used by the product.

Speaker embeddings are not secure biometric authentication and do not protect against replay or synthesized-voice attacks.

Usage

import coremltools as ct
import soundfile as sf

from frontend import compute_fbank

model = ct.models.CompiledMLModel("SpeechBrainECAPAVoxCeleb.mlmodelc")
audio, sample_rate = sf.read("voice.wav", dtype="float32")
assert sample_rate == 16000 and audio.ndim == 1
mel = compute_fbank(audio, 80)[None, :, :]
embedding = model.predict({"mel_features": mel})["embedding"]

The exact periodic Hamming window, symmetric SpeechBrain mel filters, decibel conversion, and frame limits are recorded in config.json.

Source

Converted from the official SpeechBrain checkpoint at revision 0f99f2d0ebe89ac095bcc5903c4dd8f72b367286. Checkpoint SHA-256 values and the SpeechBrain source revision used to reconstruct the graph are pinned in config.json.

Links

Downloads last month
28
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aufklarer/SpeechBrain-ECAPA-VoxCeleb-20M-CoreML

Finetuned
(15)
this model

Collection including aufklarer/SpeechBrain-ECAPA-VoxCeleb-20M-CoreML