Google launches EmbeddingGemma 2 for multimodal search
The 740 million parameter release combines text, image, video and audio embeddings for local retrieval.

Google launched EmbeddingGemma 2 on October 6, bringing text, image, video and audio retrieval into one embedding model. Code is supported within its text capability.
The Apache 2.0 release has 740 million parameters and produces vectors with 768 dimensions. Developers can load its text backbone separately from optional vision and audio encoders, reducing the components needed for a particular application.
Google’s model card describes a shared 8,192 token input budget. It warns that shrinking vectors to 128 dimensions substantially reduces multimodal quality.
Weights are available on Hugging Face. Google lists Model Garden support as coming soon.
For developers building local search, the main change is retrieving across media types with one model. Applications still need their own evaluation and safeguards.
Illustrative laptop photograph by Mohammad Rahmani on Unsplash, published December 12, 2020, under the Unsplash License. The photograph does not show EmbeddingGemma 2. Resized and converted to WebP.



