Google Open-Sources EmbeddingGemma 2: 740M Parameters Put Text, Code, Images, Video and Audio in One Vector Space, With the Text-Only Version Using About 191MB of RAM on a Phone

On October 6 Google released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license. It maps text, code, images, video and audio into a single embedding space with 740 million parameters in total, split into a 270M text backbone, a 170M vision encoder and a 300M audio encoder that can be loaded separately. Context grows from 2K tokens in the first version to 8K, enough for about 5.5 minutes of audio, 29 images or 58 video frames, and vectors can be truncated from 768 dimensions to 512, 256 or 128. Google reports 78.68 on MTEB Code, 9.92 points above the previous version; quantized, the text-only model uses about 191MB of RAM on a Pixel 11 Pro and the full multimodal model about 567MB. It is available on Hugging Face and Kaggle and supported by Ollama, llama.cpp, vLLM and others.

One model, five kinds of content

In retrieval-augmented generation (RAG), text, screenshots, meeting recordings and product videos often need separate embedding models and separate indexes that then have to be stitched together. EmbeddingGemma 2, released by Google on October 6, aims to merge that step: text, code, images, video and audio map into the same embedding space, so a single sentence can find a relevant image or audio clip.

It has 740 million parameters in total, but it is modular: a 270M text backbone, plus a 170M vision encoder and a 300M audio encoder, so text-only retrieval loads only the text part. Context grows from 2K tokens in the first version to 8K, fitting about 5.5 minutes of audio, 29 images or 58 video frames at once. Vectors are trained with Matryoshka Representation Learning and can be truncated from 768 dimensions to 512, 256 or 128, which Google says cuts storage by up to six times.

Numbers and deployment

Google reports 78.68 on MTEB Code, 9.92 points above the previous version's 68.76, and says that on MIEB (Lite) and MAEB it sets a new standard for quality per parameter among sub-1B models, outperforming some specialist models more than twice its size. On device, the quantized text-only model uses about 191MB of RAM on a Pixel 11 Pro and the full multimodal model about 567MB. It is released under Apache 2.0 on Hugging Face, Kaggle and the LiteRT Community, with support in Ollama, llama.cpp, vLLM and MediaPipe.

What it means

For developers building local knowledge bases, photo search or on-device assistants, this is a commercially usable option: under a billion parameters, able to run offline, and able to put several kinds of material in one index, saving the trouble of multiple models and alignment. Note that these results are all Google's own, and multimodal retrieval quality depends heavily on the data. Compare recall against your current text embedding setup on your own corpus first, then decide whether to move images and audio over as well.

via: Google blog