Sakura EmbeddingGemma 2 — AutoRound W4A16

Community Quantization. Not official from Google.
Quantized and evaluated by Sakura (webmp3) using Intel AutoRound optimization on top of Google's official embeddinggemma-2 architecture.


Overview

This repository provides an optimized AutoRound W4A16 (4-bit weights, 16-bit activations) quantization of google/embeddinggemma-2.

It is packaged as a standard Hugging Face repository containing packed Safetensors weights that run natively and out of the box with sentence-transformers and transformers.

  • Upstream Model: google/embeddinggemma-2
  • Pinned Upstream Commit: 914f7f89142e33e77833254d9c9b90c3cef7303b
  • Base Architecture: EmbeddingGemma2ForSequenceClassification / EmbeddingGemma2TextModel
  • Model File Size: model.safetensors: 1,296,471,072 bytes / 1,296.47 MB / 1,236.41 MiB. The pinned BF16 original model file is 1,488,915,288 bytes / 1,488.92 MB / 1,419.94 MiB: 12.93% smaller on disk. This compares model files, not total directories or measured runtime memory.
  • Verification Status: Public Hub access verified; local release smoke test passed; packaged for Transformers & SentenceTransformers.
  • Quantization Framework: AutoRound 0.16.0, as recorded in the published quantization_config.json.
  • Quantization Scheme: W4A16 symmetric (bits=4, group_size=128, iters=100); recorded packing format auto_round:auto_gptq.
  • Release Decision: RELEASE GO (Early community release based on empirical fidelity retention)

Architectural Breakdown: What is W4 vs. What Remains BF16

We explicitly do not claim "Full INT4/Q4". High-fidelity embedding models require careful treatment of sensitive components:

Component Precision Details
Text Backbone Linear Layers W4A16 216 Linear layers in language_model.layers.0 through language_model.layers.23 (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj). Group size 128, symmetric.
Token Embeddings BF16 language_model.embed_tokens (256,000 vocab) preserved in bfloat16 to avoid semantic vocabulary collapse.
Embedding Projection Head BF16 language_model.embedding_projection (768d output head) preserved in bfloat16 to preserve precise directional geometry.
Normalization Layers BF16 All RMSNorm and LayerNorm modules preserved in bfloat16 to prevent activation scale clipping.
Vision & Audio Towers BF16 Multimodal encoders (vision_tower, audio_tower) and multimodal projection heads are preserved unquantized.

Fidelity & Benchmark Results

All evaluations were conducted against an unquantized bfloat16 reference across 100 query/document retrieval pairs (English and German) and 50 code retrieval pairs.

Release Gate Status

The candidate was evaluated against strict quality criteria:

  • Empirical Gate Result: RELEASE GO / Early Community Release
  • Gate Context: The initial theoretical target of $\ge 0.99$ Mean Cosine Similarity and $\ge 95%$ Top-5 Retrieval Agreement was not fully reached in every dimension (achieved: 0.9878 Mean Cosine, 87.6% Top-5 Agreement). However, with 99.0% Top-1 Agreement, 100.0% Recall@5, and 0.9781 Spearman Correlation, these benchmark results show retained retrieval success alongside measurable BF16-fidelity loss; no NaN or Inf anomaly was observed.

The companion GGUF HQ releases achieve higher reported BF16 fidelity on the same existing BF16 reference and text/code benchmark. They are separate quantization runs; their results do not change this Safetensors model's release-gate outcome.

GGUF companion tier MiB Combined Mean Cosine (768d) Combined Spearman (768d) Text Top-5 Text/Code Recall@5
Q4 HQ 169.79 0.99400270 0.98857239 90.00% 100% / 100%
Q5 HQ — recommended balance 199.88 0.99828082 0.99664205 95.40% 100% / 100%
Q6 HQ 231.75 0.99940187 0.99879852 96.80% 100% / 100%

Values are quoted from the GGUF README's Side-by-Side Comparison. Aggregation differs: this Safetensors section reports alignment over 200 text query/document vectors; the GGUF combined Mean Cosine/Spearman include all 300 text-and-code vectors. Text Top-5 uses the 100-pair text retrieval benchmark in both. These are not identically aggregated mean scores, and no all-metric comparison with Unsloth is implied.

MRL (Matryoshka Representation Learning) Dimensions

EmbeddingGemma 2 supports dimension truncation followed by L2-renormalization. The W4A16 model demonstrates robust stability across all standard MRL truncations:

MRL Dimension Mean Cosine Sim Min Cosine Sim Spearman Rank Corr Mean L2 Drift Top-1 Agreement Top-5 Agreement Recall@5 nDCG@10
768d (Full) 0.9878 0.9671 0.9781 0.1557 99.0 % 87.6 % 100.0 % 0.9655
512d 0.9882 0.9675 0.9782 0.1531 99.0 % 88.4 % 100.0 % 0.9583
256d 0.9894 0.9691 0.9774 0.1451 100.0 % 86.2 % 100.0 % 0.9595
128d 0.9922 0.9734 0.9744 0.1246 100.0 % 84.8 % 100.0 % 0.9482

Language & Domain Subsets (768d)

Subset Mean Cosine Sim Top-1 Agreement Top-5 Agreement Recall@5
English Queries & Docs 0.9890 100.0 % 88.0 % 100.0 %
German Queries & Docs 0.9867 98.0 % 87.2 % 100.0 %
Code Retrieval 0.9859 100.0 % 88.0 % 100.0 %

Runtime & Community Format Comparison

Distribution / Format Runtime Compatibility Direct Transformers / ST Support Notes
Sakura AutoRound W4A16 (This repo) Python, PyTorch, Transformers, SentenceTransformers Yes (Plug-and-play) Runs directly in existing Python AI pipelines without needing custom binary builds.
GGUF Community Releases (unsloth, ggml-org) llama.cpp No (requires llama-server or bindings) EmbeddingGemma 2 support is upstream in llama.cpp via PR #30054, merged 2026-10-06. The earlier unknown model architecture result came from an older local build; use a build that includes this support. Community GGUF controls have since been loaded and benchmarked with a supported build.
ONNX Community Releases (onnx-community) ONNX Runtime / Transformers.js No (ONNX graph format) Modular multi-graph export tailored for WebGPU/JavaScript execution.
Sakura AutoRound GGUF HQ (Companion) llama.cpp Via llama-server / bindings Q4 HQ 169.79 MiB, Q5 HQ 199.88 MiB (recommended balance), Q6 HQ 231.75 MiB. This Safetensors/ST package is for native Python pipelines; the standalone text GGUF variants are for llama.cpp.

Safetensors vs. GGUF HQ: separate quantization runs

These are separate optimization runs, not two exports of one quantized state. The published Safetensors quantization_config.json records a different scheme and iteration count from the preserved GGUF cal256 build configurations/logs. Both use AutoRound 0.16.0 and the same base model, but that does not imply identical optimized weights or calibration.

Setting This Safetensors W4A16 release GGUF HQ cal256 runs
Transformer weight types 4-bit symmetric W4A16 Q4_K/Q5_K/Q6_K base with 48 Q8_0 block-PLE linears
Weight grouping Group size 128 Q4_K/Q5_K: 32 weights per subgroup, 8 subgroups per 256-weight superblock; Q6_K: 16, 16 subgroups per 256-weight superblock. Q8_0 uses 32-weight blocks.
Iterations 100, published configuration 50, saved build configurations/logs
Calibration selection Existing W4A16 pipeline selects up to 64 entries from calibration/calib_data.json at seqlen 128; exact historical input content is not independently established 256 synthetic retrieval samples: 48 EN queries, 48 EN docs, 48 DE queries, 48 DE docs, 32 code queries and 32 code snippets; zero exact benchmark overlap
Optimization/export path AutoRound W4A16; published packing format auto_round:auto_gptq; Safetensors for Transformers/ST Native AutoRound SignRoundV2 (enable_alg_ext=True), matching GGUF optimized-state packing, then selected embedding-path precision overrides
AutoRound version 0.16.0, published configuration 0.16.0, preserved builds
Scope Full multimodal checkpoint, vision/audio and sensitive components retained in BF16 Standalone text GGUF; CPU/native llama.cpp runtime

Historical evidence limit: the Safetensors model's local SHA256 matches the published weight file, and its local configuration matches the published configuration. The W4A16 pipeline script and saved benchmark report are available, but no immutable historical calibration-input hash/sample manifest or complete W4A16 training log was found in the inspected release records. Therefore, the exact historical 64 texts and their realized selection cannot be certified from those records alone. The script's 64-sample selection is reported as pipeline configuration, not as a recovered immutable run manifest. Exact historical Transformers/SentenceTransformers versions were not recorded in the published quantization configuration.

The GGUF cal256 runs do retain explicit sample count, calibration SHA256 b65d22bad14722ab02d316ddae2ea6f860b92c94669f5717321b66e81abb91b8, resolved per-layer configuration, native optimized-state packing evidence and training logs. These differences are enough to rule out a shared optimized quantization state. Public GGUF settings and runtime provenance are linked from the companion README; the inspected W4A16 pipeline settings are summarized here without claiming historical inputs that were not preserved.


Quickstart & Usage

1. With SentenceTransformers (Recommended)

from sentence_transformers import SentenceTransformer
import torch

# Load the quantized model
model = SentenceTransformer(
    "webmp3/Sakura-EmbeddingGemma-2-AutoRound",
    model_kwargs={"torch_dtype": torch.bfloat16}
)

# Text Retrieval Query (using official prompt_name)
query = "What is quantum entanglement?"
query_embedding = model.encode(query, prompt_name="SearchQuery")

# Documents (unprompted)
docs = [
    "Quantum entanglement is a phenomenon where particles remain connected regardless of distance.",
    "The recipe for chocolate chip cookies requires flour, butter, and sugar."
]
doc_embeddings = model.encode(docs)

# Compute similarity
similarities = model.similarity(query_embedding, doc_embeddings)
print("Similarities:", similarities)

2. Matryoshka Dimension Truncation (MRL)

To reduce memory and storage footprint, simply slice the vector and re-normalize:

import torch
import torch.nn.functional as F

# 128-dimensional embedding
full_embedding = model.encode(["Example sentence"], convert_to_tensor=True)
mrl_128 = full_embedding[:, :128]
mrl_128_normalized = F.normalize(mrl_128, p=2, dim=-1)
print("128d shape:", mrl_128_normalized.shape)

3. With Hugging Face Transformers

from transformers import AutoModel, AutoTokenizer
import torch

model_id = "webmp3/Sakura-EmbeddingGemma-2-AutoRound"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, torch_dtype=torch.bfloat16)

inputs = tokenizer(["SearchQuery: What is machine learning?"], return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

Limitations & Honest Disclosure

  1. Text Precision and File Size: The 216 text-backbone linear layers use 4-bit packed weights. model.safetensors is 1,296,471,072 bytes (1,296.47 MB / 1,236.41 MiB), compared with 1,488,915,288 bytes (1,488.92 MB / 1,419.94 MiB) for the pinned full BF16 original. The 12.93% model-file saving is modest: token embeddings and multimodal towers remain BF16. These on-disk measurements do not establish a runtime-memory reduction; the former ~5.5 GB BF16 comparison is not supported by the checked model files.
  2. Multimodal Towers: The vision and audio towers remain in 16-bit precision. If you do not use vision or audio inputs, memory consumption can be minimized by only loading the text language model.
  3. Execution Device: Optimal execution requires modern CPU (AVX-512 / VNNI) or GPU environments supporting accelerated bfloat16 and packed int4 kernels.

SHA256 Checksums

Every file in this release has been cryptographically verified:

ed5b25820992aef4a31b99a216fa369f1845f0b5001165866d82c81128c4cbc3  model.safetensors
ee5befff1a18299a2acdf77e2895cde534fcfd73d2793b7a12df644be0e6acdd  config.json
79bac98844a054c881016641b7ad2b97539c9a43c3bdd26205823e8fe1c85a15  quantization_config.json
031e56a498d33c349ab489a21885bcfe25b4fcba841149dc99e1e90d4a7c28f5  config_sentence_transformers.json
b1bcd9f2dce3ae863b359e87d0710b5dbc3314a59ecb4e2f97c7778fc8e4b228  sentence_bert_config.json
3d02572a0455b832de67fb8e63a54981bc7e8b46e337c95e917bd8122a533bfd  modules.json
ea2ae257e901064abdd98dceb19f2b0da06af600bed15e0f99f5c85c37ee9d78  preprocessor_config.json
168f6a08522f3ce5dea596d94d003af2fd691742d4f41fe1f9d8cce76bfbf69c  processor_config.json
4d777ef5bdc1aa36227abdfb77c3e49e7b9c892d16e1b6bda41c393504828be4  tokenizer.json
17bd5d6e9364ca49a534e1502076593317c298d4a663623091ed45388f004874  tokenizer_config.json
4b852efc0b9960283e735363331e6f325b33bc74bdbaa076f595bc4e9b94d85e  chat_template.jinja
8759bdf7c77efc7df7723f64856a593c8943b71ee38baf2a88771fbaf78438f9  1_Pooling/config.json
cdb09dfca347a56aa2d691744e38d5ad3c7cbc2834e7181272b9a15328b82524  2_Normalize/config.json

Citation & Acknowledgements

Downloads last month
18
Safetensors
Model size
0.6B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound

Quantized
(22)
this model