DEV Community

Prabhakar Chaudhary
Prabhakar Chaudhary

Posted on

EmbeddingGemma 2: How Google's Open Multimodal Embedding Model Works

EmbeddingGemma 2: How Google's Open Multimodal Embedding Model Works

If you've built a RAG pipeline or a semantic search system, you've almost certainly reached for a separate embedding model for each modality — one for text, another for images, maybe a third for audio. Google DeepMind's EmbeddingGemma 2, released October 6, 2026, takes a different approach: a single 740M-parameter open model that maps text, code, images, video, and audio into one shared vector space, designed to run entirely on-device.

The original EmbeddingGemma (2025) was text-only and accumulated over 20 million downloads. This release extends it to five modalities while keeping the model small enough to run on a phone. Here's what's actually inside it and what it means for practitioners.

The Architecture: One Space, Modular Encoders

EmbeddingGemma 2 is built on the Gemma 4 transformer backbone — 24 layers, grouped-query and multi-query attention, a 262,144-token vocabulary, and mean pooling with a projection layer from 512 to 768 dimensions. The key design decision is that all modalities project into the same 768-dimensional embedding space, which means a text query and an audio clip can be compared directly by cosine similarity.

The model is modular. From a single checkpoint, you load only the encoders you need:

  • Text and code only: 270M parameters (130M transformer backbone + 140M embedder)
  • Text + vision: 440M parameters (add the 170M vision encoder)
  • Text + audio: 570M parameters (add the 300M audio encoder)
  • Full multimodal: 740M parameters

This matters for deployment. If your app only does text-to-image retrieval, you're not paying the memory cost of the audio encoder.

The model shares its text tokenizer and audio encoder architecture with Gemma 4. If you're already running Gemma 4 for generation, you can pair EmbeddingGemma 2 for retrieval with a lower combined memory footprint than running two independent models.

The Token Budget: How Media Gets Encoded

All five modalities share an 8,192-token context window — four times larger than EmbeddingGemma 1. Each modality consumes that budget at fixed rates:

Modality Tokens per unit Max per pass
Images 280 per image 29 images
Video frames 140 per frame 58 frames
Audio 25 per second ~5.5 minutes
Text/code Standard tokenization Up to 8K tokens

Inputs can be interleaved — text and images mixed in one pass, with placeholder tokens marking each media item's position. This is what enables queries like "find the video clip where someone says X while showing Y on screen."

Matryoshka Representation Learning: Tunable Storage

EmbeddingGemma 2 uses Matryoshka Representation Learning (MRL), a training technique that makes the first N dimensions of any output embedding a valid, high-quality embedding on their own. At inference time, you can truncate from 768 dimensions down to 512, 256, or 128 without retraining.

The practical tradeoff, from the model card:

  • 256 dimensions: ~95% of full quality on multimodal retrieval, ~6x storage reduction
  • 128 dimensions: ~90% on text/code, but drops to ~75% on multimodal retrieval

Storing one million 768-dimensional vectors in bfloat16 takes roughly 1.5GB. At 128 dimensions, that drops to about 250MB. For large-scale vector databases where storage is a real cost, this is a meaningful lever — and you can tune it per use case without retraining.

Benchmark Results

Google reports the following on the full-precision model:

  • MTEB Code: 78.68 (up from 68.76 in EmbeddingGemma 1, a +9.92 point improvement)
  • MTEB multilingual (v2): 61.36
  • MIEB lite: 64.64
  • MMEB v2 image retrieval: 57.28
  • Visual-document retrieval: 67.84
  • Video retrieval: 50.67
  • MSEB sound retrieval: 69.54
  • MAEB audio: 49.39

The code search improvement is notable for anyone building coding agents or codebase indexing tools. The model outperforms some specialist models more than twice its size on audio and visual benchmarks — though it's worth noting these are Google's reported numbers, and independent evaluations on your specific retrieval task will tell you more.

On-Device Performance and Deployment

With quantization on a Pixel 11 Pro:

  • Text-only: ~191MB active RAM
  • Full multimodal: ~567MB active RAM

For browser deployment, the model runs via transformers.js and WebGPU. For mobile and edge, Google AI Edge's MediaPipe and LiteRT handle cross-platform deployment. The model also works with the standard server-side stack: transformers, sentence-transformers (6.1.0+), vLLM, llama.cpp, SGLang, Ollama, and LMStudio. Vector storage integrations include Qdrant.

Weights are available on Hugging Face and Kaggle under the Apache 2.0 license — commercially usable without restrictions.

Limitations to Know Before Deploying

A few things the model card is explicit about:

No safety tuning. EmbeddingGemma 2 is a pre-trained embedding model with no post-training alignment or output-level moderation. Safety mitigations are limited to pre-training data filtering (CSAM filtering, automated PII removal). If your retrieval pipeline surfaces sensitive content, you need to add your own application-level safeguards.

Use bfloat16 or float32, not float16. The model's activation range exceeds float16's dynamic range, which can produce NaN values or silently degraded embeddings. This is easy to miss if you're defaulting to float16 for efficiency.

Task instruction prefixes matter. For text inputs, the model expects task-specific instruction prefixes (e.g., "Represent this code for retrieval:"). Omitting them reduces embedding precision. The developer guide covers the recommended prefixes per task type.

MRL quality at 128 dims drops on multimodal. If you're doing cross-modal retrieval (text-to-audio, image-to-video), 128-dimensional vectors retain only ~75% of full quality. For text-only workloads, 128 dims is fine.

What This Means for Your Stack

The most immediate use case is replacing multiple single-modality embedding models with one. If you're building a RAG pipeline over a mixed corpus — documents, screenshots, audio recordings, video clips — EmbeddingGemma 2 lets you index everything into one vector database and query across modalities with a single model call.

The on-device angle is also worth taking seriously. A 191–567MB footprint means you can run retrieval entirely locally, which matters for privacy-sensitive applications (medical records, legal documents, personal media) and for offline-capable apps. Paired with Gemma 4 for generation, you get a complete on-device RAG system.

For coding agents specifically, the +9.92 MTEB Code improvement over EmbeddingGemma 1 makes this a strong candidate for codebase indexing and semantic code search — tasks where the previous version was already competitive.

The model is available now. If you're already using EmbeddingGemma 1, the upgrade path is straightforward: same ecosystem integrations, same license, substantially expanded capability.

Top comments (0)