Instructions to use srv-sngh/embeddinggemma-2-mlx-5bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use srv-sngh/embeddinggemma-2-mlx-5bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download srv-sngh/embeddinggemma-2-mlx-5bit --local-dir embeddinggemma-2-mlx-5bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
EmbeddingGemma 2, MLX 5bit
A quantized MLX build of google/embeddinggemma-2 (740M parameters; text, image, audio and video embeddings in one model.safetensors), made with Krill. Bits: 5. Weights file: 533.1 MB (bf16 original: 1488.9 MB).
All 9 builds, next to Google's original, are in the EmbeddingGemma 2 MLX collection.
What this build is
Affine 5-bit quantization, group size 64, applied to every quantizable tensor.
Produced with:
krill quantize <bf16-dir> --bits 5 --group-size 64 --dtype bf16 --output-dir <out>
Gate result: PASS (4/5-bit class). Drops against bf16: SciFact 0.41, Hindi 0.19, image R@1 0.0 points (negative means the build scored higher). Gate: drops <= 3.0 and text cosine mean >= 0.97.
Quality against bf16
Cosine similarity is measured against Google's fp32 sentence-transformers reference vectors; retrieval is nDCG@10 or Recall@k. The bf16 column is the unmodified Google weights run through Krill.
| metric | bf16 | this build |
|---|---|---|
| Text cosine to Google fp32 (min / mean) | 1.000 / 1.000 | 0.989 / 0.995 |
| Image cosine (min / mean) | 0.999 / 1.000 | 0.994 / 0.995 |
| Mixed cosine (min / mean) | 1.000 / 1.000 | 0.993 / 0.995 |
| Audio cosine (min / mean) | 1.000 / 1.000 | 0.994 / 0.995 |
| Video cosine (min / mean) | 0.999 / 1.000 | 0.995 / 0.995 |
| SciFact nDCG@10, 768 d | 86.92 | 86.51 |
| SciFact nDCG@10, 256 d | 84.44 | 83.88 |
| Hindi nDCG@10 (IndicQARetrieval) | 72.78 | 72.59 |
| Kannada nDCG@10 (IndicQARetrieval) | 73.25 | 73.25 |
| Flickr image-to-text R@1 / R@5 | 99.5 / 100.0 | 99.5 / 100.0 |
| Flickr text-to-image R@1 / R@5 | 96.1 / 99.7 | 96.1 / 99.6 |
| Clotho audio-to-text R@1 / R@5 | 25.0 / 61.7 | 26.7 / 58.3 |
| Clotho text-to-audio R@1 / R@5 | 30.0 / 58.3 | 28.3 / 58.3 |
| Ranking agreement, Spearman raw / Document | 1.000 / 1.000 | 0.992 / 0.993 |
| Top-1 neighbour unchanged, raw / Document (of 21) | 21 / 21 | 20 / 20 |
SciFact is MTEB mteb/scifact (300 queries, 5,183 documents). Hindi and Kannada are the hi and kn splits of MTEB IndicQARetrieval; the corpora are small, so treat differences under about 0.5 point as noise. Flickr is the first 200 images of the Flickr30k 1K test split. Clotho is 60 clips of the Clotho v2 test split (one clip is 1.7 points of Recall@1). Test sets are small; "PASS" means no measurable loss on these sets, not "equal".
Speed and memory (Apple M4 Pro)
| contender | weights MB | single query p50 ms | batch-32 ~256 tok docs/s | batch-32 ~1k tok docs/s | peak memory MB |
|---|---|---|---|---|---|
| Krill bf16 (fp32 compute) | 1488.9 | 10.2 | 39 | 9 | 3234 |
| Krill 5bit (this build) | 533.1 | 9.1 | 34 | 7 | 2401 |
| Google sentence-transformers, MPS bf16 | 1488.9 | 26.2 | 26 | 5 | 17283 |
Ollama embeddinggemma-2 (nvfp4, 1.3 GB with towers) |
1300 | 25.2 | 32 | 8 | 1476 |
| Unsloth Q8_0 GGUF (text only, llama.cpp Metal) | 309.9 | 16.4 | 33 | 7 | 1825 |
Unsloth and Ollama are text only; Krill rows include the vision and audio towers in the weights. Image encoding for this build: 412 ms p50 for one 640x480 PNG. Cold first request: 985 ms. Single-query p95: 12.9 ms.
Method: Apple M4 Pro, 24 GB. One contender at a time, with the server started and stopped for every repetition, 10 discarded warm-up requests, and 3 repetitions per metric reported as the median. Before each repetition the harness required 1-minute load below 3, at least 50% free memory, no thermal limit and AC power. A contender counts as stable when the spread (max - min over the median) of every metric across its 3 repetitions is at most 10%; otherwise it was re-measured. Krill rows use fp32 compute (the default). Full protocol: docs/bench/embeddinggemma2-2026-10-08.md in the Krill repository.
Usage with Krill
krill pull embeddinggemma-2-5bit
krill serve
curl localhost:57455/v1/embeddings -H 'Content-Type: application/json' \
-H "Authorization: Bearer $KRILL_API_KEY" -d '{"model":"embeddinggemma-2-5bit","input":"What causes the northern lights?","task":"SearchQuery"}'
Use "task": "SearchQuery" for queries and "Document" for documents. Dimensions: 768 (default), 512, 256, 128 (Matryoshka). Context: 8,192 tokens. Compute dtype: float32 (default) or bfloat16; float16 is not supported.
Other builds
mixed-mxfp8, 8bit, 6bit, 6bit-dyn, 5bit, 4bit-g32, 4bit-dyn, 4bit-dyn-text and nvfp4, all under srv-sngh/embeddinggemma-2-mlx-<name>.
- Downloads last month
- 17
5-bit
Model tree for srv-sngh/embeddinggemma-2-mlx-5bit
Base model
google/embeddinggemma-2