EmbeddingGemma 2, MLX 5bit

A quantized MLX build of google/embeddinggemma-2 (740M parameters; text, image, audio and video embeddings in one model.safetensors), made with Krill. Bits: 5. Weights file: 533.1 MB (bf16 original: 1488.9 MB).

All 9 builds, next to Google's original, are in the EmbeddingGemma 2 MLX collection.

What this build is

Affine 5-bit quantization, group size 64, applied to every quantizable tensor.

Produced with:

krill quantize <bf16-dir> --bits 5 --group-size 64 --dtype bf16 --output-dir <out>

Gate result: PASS (4/5-bit class). Drops against bf16: SciFact 0.41, Hindi 0.19, image R@1 0.0 points (negative means the build scored higher). Gate: drops <= 3.0 and text cosine mean >= 0.97.

Quality against bf16

Cosine similarity is measured against Google's fp32 sentence-transformers reference vectors; retrieval is nDCG@10 or Recall@k. The bf16 column is the unmodified Google weights run through Krill.

metric bf16 this build
Text cosine to Google fp32 (min / mean) 1.000 / 1.000 0.989 / 0.995
Image cosine (min / mean) 0.999 / 1.000 0.994 / 0.995
Mixed cosine (min / mean) 1.000 / 1.000 0.993 / 0.995
Audio cosine (min / mean) 1.000 / 1.000 0.994 / 0.995
Video cosine (min / mean) 0.999 / 1.000 0.995 / 0.995
SciFact nDCG@10, 768 d 86.92 86.51
SciFact nDCG@10, 256 d 84.44 83.88
Hindi nDCG@10 (IndicQARetrieval) 72.78 72.59
Kannada nDCG@10 (IndicQARetrieval) 73.25 73.25
Flickr image-to-text R@1 / R@5 99.5 / 100.0 99.5 / 100.0
Flickr text-to-image R@1 / R@5 96.1 / 99.7 96.1 / 99.6
Clotho audio-to-text R@1 / R@5 25.0 / 61.7 26.7 / 58.3
Clotho text-to-audio R@1 / R@5 30.0 / 58.3 28.3 / 58.3
Ranking agreement, Spearman raw / Document 1.000 / 1.000 0.992 / 0.993
Top-1 neighbour unchanged, raw / Document (of 21) 21 / 21 20 / 20

SciFact is MTEB mteb/scifact (300 queries, 5,183 documents). Hindi and Kannada are the hi and kn splits of MTEB IndicQARetrieval; the corpora are small, so treat differences under about 0.5 point as noise. Flickr is the first 200 images of the Flickr30k 1K test split. Clotho is 60 clips of the Clotho v2 test split (one clip is 1.7 points of Recall@1). Test sets are small; "PASS" means no measurable loss on these sets, not "equal".

Speed and memory (Apple M4 Pro)

contender weights MB single query p50 ms batch-32 ~256 tok docs/s batch-32 ~1k tok docs/s peak memory MB
Krill bf16 (fp32 compute) 1488.9 10.2 39 9 3234
Krill 5bit (this build) 533.1 9.1 34 7 2401
Google sentence-transformers, MPS bf16 1488.9 26.2 26 5 17283
Ollama embeddinggemma-2 (nvfp4, 1.3 GB with towers) 1300 25.2 32 8 1476
Unsloth Q8_0 GGUF (text only, llama.cpp Metal) 309.9 16.4 33 7 1825

Unsloth and Ollama are text only; Krill rows include the vision and audio towers in the weights. Image encoding for this build: 412 ms p50 for one 640x480 PNG. Cold first request: 985 ms. Single-query p95: 12.9 ms.

Method: Apple M4 Pro, 24 GB. One contender at a time, with the server started and stopped for every repetition, 10 discarded warm-up requests, and 3 repetitions per metric reported as the median. Before each repetition the harness required 1-minute load below 3, at least 50% free memory, no thermal limit and AC power. A contender counts as stable when the spread (max - min over the median) of every metric across its 3 repetitions is at most 10%; otherwise it was re-measured. Krill rows use fp32 compute (the default). Full protocol: docs/bench/embeddinggemma2-2026-10-08.md in the Krill repository.

Usage with Krill

krill pull embeddinggemma-2-5bit
krill serve
curl localhost:57455/v1/embeddings -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $KRILL_API_KEY" -d '{"model":"embeddinggemma-2-5bit","input":"What causes the northern lights?","task":"SearchQuery"}'

Use "task": "SearchQuery" for queries and "Document" for documents. Dimensions: 768 (default), 512, 256, 128 (Matryoshka). Context: 8,192 tokens. Compute dtype: float32 (default) or bfloat16; float16 is not supported.

Other builds

mixed-mxfp8, 8bit, 6bit, 6bit-dyn, 5bit, 4bit-g32, 4bit-dyn, 4bit-dyn-text and nvfp4, all under srv-sngh/embeddinggemma-2-mlx-<name>.

Downloads last month
17
Safetensors
Model size
0.7B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for srv-sngh/embeddinggemma-2-mlx-5bit

Quantized
(57)
this model

Collection including srv-sngh/embeddinggemma-2-mlx-5bit