Instructions to use aufklarer/SpeechBrain-ECAPA-VoxCeleb-20M-CoreML with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- speechbrain
How to use aufklarer/SpeechBrain-ECAPA-VoxCeleb-20M-CoreML with speechbrain:
# interface not specified in config.json
- Notebooks
- Google Colab
- Kaggle
SpeechBrain ECAPA VoxCeleb Core ML
This is a reproducible compiled Core ML export of
speechbrain/spkrec-ecapa-voxceleb.
It produces a local voice embedding for cosine-similarity speaker comparison.
Model
| Property | Value |
|---|---|
| Parameters | 20.77 million |
| Format | Compiled Core ML, FLOAT16 compute |
| Compiled size | 39.8 MiB |
| Input | SpeechBrain log-mel [1, frames, 80] |
| Sample rate | 16 kHz |
| Output | 192-dimensional L2-normalized embedding |
| Mel frames | 10 to 3,001 (about 0.1 to 30 seconds) |
| Minimum deployment | macOS 15 / iOS 18 |
The graph includes sentence mean normalization. Audio-to-mel processing stays outside the graph so apps can share one tested streaming frontend between Core ML and MLX.
Files
| File | Description |
|---|---|
SpeechBrainECAPAVoxCeleb.mlmodelc/ |
Precompiled Core ML model |
frontend.py |
Reproducible SpeechBrain log-mel frontend |
config.json |
Graph, audio frontend, source revision, and checksums |
artifact_manifest.json |
SHA-256 and size of every compiled artifact file |
validation.json |
Conversion parity and local latency measurements |
requirements.txt |
Pinned standalone runtime dependencies |
LICENSE |
Apache 2.0 license |
Validation
| Check | Result | Meaning |
|---|---|---|
| Minimum PyTorch/Core ML output cosine | 0.999995 | 1.0 is identical direction |
| Maximum absolute error | 0.000702 | Lower is better |
| Warm local inference | 5.0 ms | Median after warm-up on the conversion machine |
The official source reports 0.80% equal-error rate on the cleaned VoxCeleb1 test set. That upstream result is not presented as a new benchmark for this conversion. Before changing a speaker-verification threshold, validate the compiled artifact on VoxCeleb1-O and the microphones and languages used by the product.
Speaker embeddings are not secure biometric authentication and do not protect against replay or synthesized-voice attacks.
Usage
import coremltools as ct
import soundfile as sf
from frontend import compute_fbank
model = ct.models.CompiledMLModel("SpeechBrainECAPAVoxCeleb.mlmodelc")
audio, sample_rate = sf.read("voice.wav", dtype="float32")
assert sample_rate == 16000 and audio.ndim == 1
mel = compute_fbank(audio, 80)[None, :, :]
embedding = model.predict({"mel_features": mel})["embedding"]
The exact periodic Hamming window, symmetric SpeechBrain mel filters, decibel
conversion, and frame limits are recorded in config.json.
Source
Converted from the official SpeechBrain checkpoint at revision
0f99f2d0ebe89ac095bcc5903c4dd8f72b367286. Checkpoint SHA-256 values and the SpeechBrain source
revision used to reconstruct the graph are pinned in config.json.
Links
- speech-swift — Apple SDK
- Docs — install and CLI docs
- soniqo.audio
- blog
- Downloads last month
- 28
Model tree for aufklarer/SpeechBrain-ECAPA-VoxCeleb-20M-CoreML
Base model
speechbrain/spkrec-ecapa-voxceleb