Attributed compile + one small real run
Primary sources (all Google, 2026-10-06): EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings by Sahil Dua and Henrique Schechter Vera (Google DeepMind); EmbeddingGemma 2: The Developer Guide by Maarten Grootendorst and Ian Ballantyne; Bring multimodal semantic search to the edge with EmbeddingGemma 2 (Google AI Edge and Android ML teams); the google/embeddinggemma-2 model card on Hugging Face
Secondary coverage: Google releases a 740M-parameter embedding model for local multimodal search by Ryan Merket (Runtime Wire, 2026-10-06)
Model specs, benchmark scores, and on-device RAM figures below are Google's own numbers. The only thing I measured is the tiny text-only code-search test in this post, run once on a CPU-only Linux box. It is a smoke test, not a benchmark. I did not test images, video, or audio.
Google released EmbeddingGemma 2 yesterday: an Apache 2.0 embedding model, built on Gemma 4, that puts text, code, images, video, and audio into one 768-dimensional vector space. The first EmbeddingGemma was text-only and, per Google, passed 20 million downloads. This one is the "search my photos with a sentence, without uploading them" version.
Most coverage focuses on the multimodal demos. As a backend engineer, I care about a narrower question: can I drop the text-and-code part into a RAG pipeline today, on a CPU, and what bites first? I ran it to find out.
What shipped (vendor-reported)
- Modular size. 740M parameters in total, but you load only what you need. Text and code alone is 270M. Adding vision makes it 440M, adding audio instead makes it 570M, and everything together is 740M. Google's AI Edge team reports about 191 MB of active RAM for text-only and about 567 MB for the full model on a Pixel 11 Pro.
- One shared space. A text query can be compared directly with an image, a video frame, or an audio clip. All inputs share an 8,192-token budget. Google prices an image at 280 tokens, a video frame at 140 tokens (sampled at 1 fps by default), and audio at 25 tokens per second.
- Matryoshka truncation. You can keep only the first 512, 256, or 128 dimensions of each vector to save storage. Google's table shows the cost: the code benchmark drops from 78.68 at 768d to 71.41 at 128d, and the overall multimodal score drops from 59.01 to 45.65. Google recommends 128d mainly for text-only workloads.
- Code is the headline text gain. On MTEB Code, Google lists 78.68 versus 68.76 for the first EmbeddingGemma. Multilingual MTEB barely moved (61.36 versus 61.15). So if you only embed prose, the upgrade is mostly about multimodality, not text quality.
-
Task prefixes matter. The model was trained with short instructions such as
task: code retrieval | query: ...for queries andtitle: "... | text: ...for documents. Sentence Transformers exposes these as named prompts."
The run: natural-language code search on a CPU
My setup: a CPU-only Linux box with 8 vCPUs, Python 3.13, torch 2.14.1+cpu, sentence-transformers 6.1.0, and transformers 5.19.0. The weights are not gated on Hugging Face, so there is no license click-through or token.
I wrote eight small Python snippets (retry with backoff, slugify, JWT signature check, text chunker, token-bucket rate limiter, CSV dedupe, a FastAPI health check, cosine similarity). Then I asked six plain-English questions that deliberately avoid the function names, like "back off and try again when a call fails."
import torch
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None}, # text + code only
model_kwargs={"torch_dtype": torch.float32}, # never float16
)
docs = [f"title: "{name} | text: {code}\" for name, code in snippets.items()]"
d = model.encode(docs, truncate_dim=768, normalize_embeddings=True)
q = model.encode(queries, prompt_name="CodeRetrieval",
truncate_dim=768, normalize_embeddings=True)
scores = q @ d.T # cosine similarity, since both sides are unit length
What happened on this run:
- The text-only model loaded in 6.4 seconds and reported 271M parameters, which matches Google's 270M text configuration.
-
All six queries put the right snippet first, at both 768d and 128d. For example, "check that a signed auth token wasn't tampered with" returned
jwt_check.py, and "remove duplicate customers from a spreadsheet export" returnedcsv_dedupe.py. - Encoding the eight snippets plus six queries took about half a second on the CPU.
Eight documents is far too few to say anything about quality at scale. What it does show is that the text path runs comfortably on an ordinary CPU, and that the CodeRetrieval prompt works as documented.
Gotcha 1: "text-only" still needs the image stack to load
My first two attempts crashed before any embedding happened. Even with vision_config=None and audio_config=None, loading failed first with EmbeddingGemma2Processor requires the PIL library, and after installing Pillow, with No module named 'torchvision'. In this version combination the processor imports Gemma 4's image-processing code no matter which encoders you enable.
The fix that worked for me:
pip install -U sentence-transformers transformers pillow
pip install --index-url https://download.pytorch.org/whl/cpu torch torchvision
If you are building a slim Docker image for a text-only RAG service, budget for both packages, or pin versions and check again when a later release decouples them.
Gotcha 2: slicing vectors without re-normalizing
Matryoshka truncation tempts you to store 768d vectors and slice them later. The model card warns that sliced vectors are no longer unit length, and that cosine scores then degrade "silently." I checked what that looks like in practice.
After slicing my normalized 768d document vectors to their first 128 values, their lengths were about 0.58 to 0.60 instead of 1.0. The score for the retry query against retry.py fell from 0.849 (properly re-normalized 128d) to 0.295 (raw slice). Nothing errors. If your retriever has a similarity cutoff such as "drop anything under 0.5," the raw slice silently returns nothing.
The safe version is to let the library do it:
model.encode(texts, truncate_dim=128, normalize_embeddings=True)
Also keep query and document dimensions identical. A 768d query can't be scored against a 128d index.
A related surprise from the same run: the correct 128d scores were all higher than the 768d ones (0.82 to 0.87 versus 0.75 to 0.79). The rankings matched, but the absolute numbers moved. Any threshold you tuned at one dimension needs re-tuning at another.
Gotcha 3: float16 (per Google, not tested)
Google's guide says the model's activations overflow float16, which produces NaN or degraded vectors without raising an error. Use bfloat16 on GPUs that support it and float32 on CPUs. I only ran float32, so I'm passing this along rather than confirming it. It matters because "cast to fp16 to save memory" is a common reflex.
A small prompt-name note
Google's current docs use names like SearchQuery, Document, and CodeRetrieval. The checkpoint I loaded also still exposes the older names (Retrieval-query, Retrieval-document, STS) from EmbeddingGemma 1. Pick one convention and use it for both indexing and querying. Mixing a query prefix from one family with documents embedded under another is the kind of mismatch that never throws an error.
When to reach for it
- Private, local search over notes, code, screenshots, or recordings that shouldn't leave the machine. This is the case Google is designing for.
- Code search inside a dev tool, where the 270M text configuration is small enough to ship with the tool.
- Zero-shot intent routing. Google's edge post pitches matching user input directly against label descriptions, with no training data, as a cheap router in front of a bigger model.
If you already pay for a hosted embedding API and your data can leave the device, Google's own cloud counterpart, Gemini Embedding 2, is the comparison to run on your own corpus before switching.
Checklist
- Install
pillowandtorchvisioneven for text-only use. - Load only the encoders you need via
config_kwargs. - Use
float32on CPU andbfloat16on GPU. Avoidfloat16. - Use the right task prefix:
CodeRetrievalorSearchQueryfor queries, andtitle: "... | text: ...for documents." - Truncate with
truncate_dimplusnormalize_embeddings=True, never by slicing. - Re-tune similarity thresholds whenever you change dimensions.
- Validate on 50 to 100 of your own queries before trusting the 128d setting.
Byline: YongBo Yu — Toronto AI engineer (agents, LLM workflows, RAG). GitHub: YongBoYu1.
Top comments (0)