DEV Community

Cover image for EmbeddingGemma 2 puts multimodal search on a phone
techaiwire
techaiwire

Posted on Originally published at techaiwire.com

EmbeddingGemma 2 puts multimodal search on a phone

Google DeepMind released EmbeddingGemma 2 on October 6, 2026, an open model that turns text, code, images, video and audio into vectors in one shared space. It has 740 million parameters and ships under the Apache 2.0 license, according to Google's developer guide. The text-only version runs in about 191 MB of memory on a phone. That makes it practical to search a user's photos, voice notes and documents without sending them to a server.

What an embedding model does

An embedding model turns a piece of content into a list of numbers, called a vector. Similar content ends up with similar vectors. Search then becomes a matter of finding the nearest vectors to a query.

This is the engine behind RAG, or retrieval-augmented generation, where an app finds relevant documents before asking a language model to answer. EmbeddingGemma 2 puts every kind of media in the same space. A text query like "dog on a beach" can therefore find a photo, a video frame or an audio clip directly.

The edge guide describes it as a model that "maps text, images, video frames, and audio into a single, unified vector space."

How it is built

The model is made of stackable encoders that share one embedding space:

Part Parameters
Text and code encoder 270M
Vision encoder 170M
Audio encoder 300M
Total 740M

Stackable means a text-only app can load just the 270M core, and add vision or audio only if it needs them.

The output is a 768-dimension vector. It can be cut down to 512, 256 or 128 dimensions using Matryoshka Representation Learning, a training method that packs the most important information into the front of each vector. Google says 128 dimensions cuts storage by 6x. "Most of the full quality of the original embedding on text and code is retained and about 95% on image, video, and speech retrieval," the developer guide says.

The context window is 8,192 tokens, four times the first version's. Unite.AI lists what fits in one call:

  • 29 images, at 280 tokens each
  • 58 video frames, at 140 tokens each
  • 5.5 minutes of audio, at 25 tokens per second

Unite.AI adds that the model supports more than 100 languages and has a training data cutoff of January 2025.

Benchmarks and speed

Test EmbeddingGemma 2 Source
MTEB Code 78.68 (EmbeddingGemma 1: 68.76) Unite.AI, Crypto Briefing
MTEB multilingual 61.36 Unite.AI
MMEB v2 image retrieval 57.28 Unite.AI
MMEB v2 video retrieval 50.67 Unite.AI
MSEB audio 69.54 Unite.AI

The code score is the biggest jump, a gain of 9.92 points. Google's guide describes it as 14% higher than the first version.

On hardware, Google reports about 191 MB of RAM for text-only use and about 567 MB for full multimodal use on a Pixel 11 Pro, with quantization. The edge guide measured image embedding at 37.3 milliseconds on a MacBook M5 Pro GPU, or 26.9 images a second.

The same guide says the model can sort an input among 500 options in under 100 milliseconds, with no training needed. That puts it close to the new class of decision models such as Cloudflare's Clef.

Where to get it

The weights are on Hugging Face as google/embeddinggemma-2, and on Kaggle and the LiteRT Community. Crypto Briefing lists support in sentence-transformers, Hugging Face transformers, LiteRT and MediaPipe, MLX and Ollama. Google also ships pre-quantized LiteRT bundles as .litertlm files.

What this means for developers

  • Move search onto the device. At about 191 MB for text, the model fits in a mobile app. Private content can be indexed locally, which Unite.AI notes means personal files "don't have to be uploaded to a server just to become searchable."
  • Use truncation to cut storage. If your vector database is large, test 256 or 128 dimensions. Google reports a 6x saving at 128 with most text quality kept.
  • Re-embed if you change models. Vectors from EmbeddingGemma 1 and 2 are not interchangeable. Plan a full re-index rather than mixing them.
  • Try it for code search. A 9.92-point gain on MTEB Code makes it worth testing for repository search and code RAG.
  • Load only the encoders you need. A text-only feature does not need the 470M parameters of vision and audio encoders. Smaller loads mean faster start-up on phones.

Run it against your current embedding model on your own queries first. The public benchmarks show the direction, and your own data shows the real gain.


This article was first published on Tech AI Wire.

Also available in

Deutsch · 日本語 · Français · Español · Português

Related on Tech AI Wire

Sources

Top comments (0)