One model, five kinds of content
In retrieval-augmented generation (RAG), text, screenshots, meeting recordings and product videos often need separate embedding models and separate indexes that then have to be stitched together. EmbeddingGemma 2, released by Google on October 6, aims to merge that step: text, code, images, video and audio map into the same embedding space, so a single sentence can find a relevant image or audio clip.
It has 740 million parameters in total, but it is modular: a 270M text backbone, plus a 170M vision encoder and a 300M audio encoder, so text-only retrieval loads only the text part. Context grows from 2K tokens in the first version to 8K, fitting about 5.5 minutes of audio, 29 images or 58 video frames at once. Vectors are trained with Matryoshka Representation Learning and can be truncated from 768 dimensions to 512, 256 or 128, which Google says cuts storage by up to six times.
Numbers and deployment
Google reports 78.68 on MTEB Code, 9.92 points above the previous version's 68.76, and says that on MIEB (Lite) and MAEB it sets a new standard for quality per parameter among sub-1B models, outperforming some specialist models more than twice its size. On device, the quantized text-only model uses about 191MB of RAM on a Pixel 11 Pro and the full multimodal model about 567MB. It is released under Apache 2.0 on Hugging Face, Kaggle and the LiteRT Community, with support in Ollama, llama.cpp, vLLM and MediaPipe.
What it means
For developers building local knowledge bases, photo search or on-device assistants, this is a commercially usable option: under a billion parameters, able to run offline, and able to put several kinds of material in one index, saving the trouble of multiple models and alignment. Note that these results are all Google's own, and multimodal retrieval quality depends heavily on the data. Compare recall against your current text embedding setup on your own corpus first, then decide whether to move images and audio over as well.
via: Google blog