LiteRT-LM

litert-community/embeddinggemma-2-740m-litert-lm

Main Model Card: google/embeddinggemma-2

EmbeddingGemma 2 740M is the full, natively multimodal variant of the EmbeddingGemma 2 family, packaged for low-latency, on-device inference using LiteRT. Integrating the 270M parameter text backbone with modular vision (170M) and audio (300M) encoders, this 740M parameter model maps text, code, images, video, and audio into a single, shared 768-dimensional vector space. It enables robust cross-modal semantics and supports interleaved multimodal inputs within a single 8,192-token context window. Built for ultimate flexibility, it allows developers to run full multimodal search, speech retrieval, and complex cross-media semantic similarity pipelines directly on consumer hardware.

Try EmbeddingGemma 2 with LiteRT-LM

Ready to integrate this into your product? Get started here.

Try EmbeddingGemma 2 with MediaPipe

MediaPipe uses EmbeddingGemma 2 and LiteRT-LM to enable high-level, cross-platform tasks which you can experience firsthand with the following interactive web demos:

Model Specifications

EmbeddingGemma 2
Text 270M
EmbeddingGemma 2
Text Vision 440M
EmbeddingGemma 2
740M
Supported Modalities Text Text, Images Text, Images, Video, Audio
Parameters Total: 270M
Transformer: 130M
Embeddings: 140M
Total: 440M
Transformer: 130M
Embeddings: 140M
Vision Encoder: 170M
Total: 740M
Transformer: 130M
Embeddings: 140M
Vision Encoder: 170M
Audio Encoder: 300M
Quantization Scheme Transformer:
    int4 per-channel (QAT)
Embeddings:
    int4 per-channel (QAT)
Transformer:
    int4 per-channel (QAT)
Embeddings:
    int4 per-channel (QAT)
Vision Encoder:
    int8 per-channel (QAT)
Transformer:
    int4 per-channel (QAT)
Embeddings:
    int4 per-channel (QAT)
Vision Encoder:
    int8 per-channel (QAT)
Audio Encoder:
    mixed int2/int4/int8
    per-channel (QAT)
Download Size (CPU/GPU file) 165 MB 388 MB 485 MB
On-Demand Modality Loading - Yes Yes
Supported Input Sizes Text tokens: 128, 256, 512,
    1024, 2048, 8192
Text tokens: 128, 256, 512,
    1024, 2048, 8192
Image soft tokens: 70, 140
Text tokens: 128, 256, 512,
    1024, 2048, 8192
Image soft tokens: 70, 140
Audio soft tokens: 12
Supported Output Sizes 128 (512 bytes)
256 (1024 bytes)
512 (2048 bytes)
768 (3072 bytes)
128 (512 bytes)
256 (1024 bytes)
512 (2048 bytes)
768 (3072 bytes)
128 (512 bytes)
256 (1024 bytes)
512 (2048 bytes)
768 (3072 bytes)
Link litert-community/embeddinggemma-2-text-270m-litert-lm litert-community/embeddinggemma-2-text-vision-440m-litert-lm litert-community/embeddinggemma-2-740m-litert-lm

Additional Notes:

  • LiteRT-LM automatically scales and patchifies images of any size to fit EmbeddingGemma 2's vision encoder model.
  • LiteRT-LM supports EmbeddingGemma 2's streamed audio allowing a variety of different audio lengths.

EmbeddingGemma 2 Performance on LiteRT-LM

The text performance was measured by loading and running the 128 text signature. For vision benchmarking, the vision encoder used the 70 signature. The audio file for the Text + Audio run was 10 seconds long while the file for Text + Audio + Vision was 5 seconds. The latency reported is the average of 5 iterations.

Memory was measured with each platform's native metrics and is not directly comparable across operating systems. CPU memory was measured using, rusage::ru_maxrss on Android, Linux and IOT, task_vm_info::phys_footprint on iOS and MacBook, and process_memory_counters::PrivateUsage on Windows. With the exception of task_vm_info::phys_footprint, accelerator (GPU/NPU/TPU) memory is not included. Memory measurements were taken from the second load which reads from loading caches. Memory usage on the first load may vary.

Android

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
Google Pixel 11 Pro TPU 8.3 ms 49 ms 193.5 ms 149 ms
S26 Ultra CPU 27.1 ms 175 ms 305 ms 322 ms
S26 Ultra GPU 25.9 ms 119 ms 350 ms 324 ms
Device Backend Text CPU memory Text + Vision CPU memory Text + Audio CPU memory Text + Vision + Audio CPU memory
Google Pixel 11 Pro TPU 112 MB 125 MB 121 MB 127.5 MB
S26 Ultra CPU 334 MB 694 MB 465 MB 811 MB
S26 Ultra GPU 333 MB 427 MB 407 MB 493 MB

iOS

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
iPhone 18 Pro CPU 41.8 ms 191 ms 445 ms 400 ms
iPhone 18 Pro GPU 11.6 ms 69.8 ms 165 ms 150 ms
Device Backend Text CPU/GPU memory Text + Vision CPU/GPU memory Text + Audio CPU/GPU memory Text + Vision + Audio CPU/GPU memory
iPhone 18 Pro CPU 84 MB 196 MB 98 MB 196 MB
iPhone 18 Pro GPU 85 MB 196 MB 96 MB 226 MB

Linux

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
Arm 2.3 & 2.8GHz CPU 105 ms 825 ms 947 ms 1251 ms
NVIDIA GeForce RTX 4090 GPU 7.6 ms 23.9 ms 169 ms 112 ms
Device Backend Text CPU memory Text + Vision CPU memory Text + Audio CPU memory Text + Vision + Audio CPU memory
Arm 2.3 & 2.8GHz CPU 310 MB 647 MB 431 MB 764 MB
NVIDIA GeForce RTX 4090 GPU 528 MB 678 MB 654 MB 811 MB

macOS

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
MacBook Pro M5 CPU 31.3 ms 151 ms 314 ms 432 ms
MacBook Pro M5 GPU 9.5 ms 37.3 ms 195 ms 131 ms
Device Backend Text CPU/GPU memory Text + Vision CPU/GPU memory Text + Audio CPU/GPU memory Text + Vision + Audio CPU/GPU memory
MacBook Pro M5 CPU 165 MB 310 MB 201 MB 336 MB
MacBook Pro M5 GPU 233 MB 403 MB 320 MB 478 MB

Windows

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
Dell XPS 16 (Intel Core Ultra Series 3) CPU 71.7 ms 338 ms 813 ms 708 ms
Dell XPS 16 (Intel Core Ultra Series 3) GPU 19.2 ms 62.5 ms 284 ms 205 ms
Dell XPS 16 (Intel Core Ultra Series 3) Running Intel OpenVINO NPU 13.3 ms 49.8 ms 124.3 ms 108.6 ms
Device Backend Text CPU memory Text + Vision CPU memory Text + Audio CPU memory Text + Vision + Audio CPU memory
Dell XPS 16 (Intel Core Ultra Series 3) CPU 211 MB 342 MB 242 MB 369 MB
Dell XPS 16 (Intel Core Ultra Series 3) GPU 822 MB 1769 MB 1047 MB 1962 MB
Dell XPS 16 (Intel Core Ultra Series 3) Running Intel OpenVINO NPU 201 MB 394 MB 288 MB 503 MB

Web

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
MacBook Pro M5 GPU 21.8 ms 107 ms 228 ms 226 ms

IoT

Device Backend Text latency              Text + Vision latency              Text + Audio latency              Text + Vision + Audio latency             
Raspberry Pi 5 16GB CPU 161 ms 1761 ms 946 ms 2304 ms
Jetson Orin Nano CPU 207 ms 1741 ms 1659 ms 2507 ms
Jetson Orin Nano GPU 88 ms 485 ms 1195 ms 1069 ms
Arduino VENTUNO Q NPU 13.6 ms 135 ms - * - *
Device Backend Text CPU memory Text + Vision CPU memory Text + Audio CPU memory Text + Vision + Audio CPU memory
Raspberry Pi 5 16GB CPU 282 MB 619 MB 408 MB 736 MB
Jetson Orin Nano CPU 302 MB 627 MB 417 MB 741 MB
Jetson Orin Nano GPU 523 MB 1103 MB 755 MB 1228 MB
Arduino VENTUNO Q NPU 206 MB 460 MB - * - *
* Audio modality is not supported on the IQ-8275 NPU so this metric is omited.
Downloads last month
8,037
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for litert-community/embeddinggemma-2-740m-litert-lm

Finetuned
(22)
this model

Collection including litert-community/embeddinggemma-2-740m-litert-lm