Instructions to use onnx-community/embeddinggemma-300m-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use onnx-community/embeddinggemma-300m-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('sentence-similarity', 'onnx-community/embeddinggemma-300m-ONNX');
Deployment benchmark on Android / Snapdragon 8 Gen 3 (Termux + ONNX-CPU) — latency, memory, retrieval-quality A/B vs gemini-embedding-2, matryoshka, OOC calibration
Hi all — I ran a deployment benchmark of this exact ONNX model (onnx/model_quantized.onnx, INT8) on a Snapdragon 8 Gen 3 / Android 14 / Termux stack with stock ONNX Runtime CPU. As far as I could find, there was no public benchmark covering both inference characteristics and retrieval-quality drift on real Android hardware against a real RAG corpus — so I produced one.
Headline numbers
| Metric | Value |
|---|---|
| Cold start (tokenizer + session + first inference, cold-RAM) | 3.1 s — 65% of which is tokenizer load, not the model |
| Warm latency, 128-token query (n=30) | 374 ms |
| Throughput peak | 417 tok/s @ 512 tokens — drops past that as quadratic attention dominates |
| Peak RSS | 1.82 GB — 6× the on-disk model size; phones with 4 GB RAM will see memory pressure |
| Top-1 domain match vs gemini-embedding-2 (same 543-chunk corpus, pure cosine, 18 in-corpus probes) | 94% / 94% |
| Matryoshka 768d → 512d (slice + L2-renorm) | 86% chunk-identity, 94% domain correct (= 768d), 33% storage cut |
| OOC threshold recalibration | 0.35 → 0.42 is strict Pareto improvement — catches 25% more OOC at zero in-corpus loss |
What's notable
- EmbeddingGemma matches gemini-embedding-2 at the domain level (94% / 94%) on this corpus, despite the 4× smaller embedding (768d vs 3072d) and zero recurring API cost. Chunk-level disagreement (45%) is mostly "different relevant chunks of the same correct doc," not retrieval failure — Jaccard@3 of 0.53 confirms substantial top-N overlap.
- The smaller embedder is actually more OOC-separable on this corpus. EmbeddingGemma's in-corpus cosines span 0.446–0.690; OOC scores at 0.000, 0.311, 0.415, 0.510 mostly fall below the in-corpus min. gemini-embedding-2's OOC scores (0.602, 0.629, 0.679) sit inside its in-corpus range (0.604–0.806), so it can't be cleanly thresholded.
- Tokenizer load dominates cold start (2034 ms / 65%), not the model (616 ms / 20%) or first inference (453 ms / 15%). Cheap UX win: preload
AutoTokenizer.from_pretrainedon a background thread at app launch.
Repo + reproducibility
- Repo (Apache 2.0): https://github.com/verbalogicproject-creator/Edge-Benchmark-workstation
- Full report: https://github.com/verbalogicproject-creator/Edge-Benchmark-workstation/blob/main/models/embeddinggemma-300m-int8/REPORT.md
- Methodology (warm-up policy, percentile choice, prefix conventions, OOC threshold definitions): https://github.com/verbalogicproject-creator/Edge-Benchmark-workstation/blob/main/docs/methodology-embedders.md
Replication is one pip install -r requirements.txt + setting EMBEDDINGGEMMA_MODEL_PATH away. Tier 2 (quality A/B) requires GEMINI_API_KEY (~$0.16 first run, cached after). The probes set is in probes.json at the repo root — 22 probes, 6 categories, exact-domain labels.
Caveats
- Single-device result (ZTE nubia RedMagic 10 Air). Transfers roughly to peer flagship-class arm64 Androids, less so to entry-level silicon.
- CPUExecutionProvider only. No QNN / Hexagon NPU, no Adreno GPU. That's deliberately scoped as v2 work; the report's "What this did not measure" section is explicit.
- Tier 2 numbers are corpus-specific (one 543-chunk RAG corpus). The shape of the findings (domain match rate, cosine distribution, OOC overlap) generalizes; absolute numbers won't transfer to other corpora unchanged.
Happy to take questions, or hear about discrepancies on other Snapdragon-class devices. If you replicate, the methodology doc has the exact protocol.