Google released EmbeddingGemma 2 on October 6, 2026, as the successor to last year's EmbeddingGemma, which Google says passed 20 million downloads and was widely used for on-device search and privacy-first retrieval augmented generation. The new version expands beyond text to unify code, images, video and audio in a single embedding space, so a voice memo can find a video clip or a text query can search hours of audio recordings.
EmbeddingGemma 2 is built on the Gemma 4 architecture, released under the commercially permissive Apache 2.0 license, and has 740 million parameters. It is modular: text-only workloads need as little as 270M parameters, with optional vision (170M) and audio (300M) encoders for full multimodal support. Output vectors default to 768 dimensions and can be truncated to 512, 256 or 128, which Google says cuts local vector storage by up to six times.
According to Google, the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. It has an 8K token context window, four times the first version, enough for about 5.5 minutes of audio, 29 images or 58 video frames. Code performance on MTEB Code rose from 68.76 to 78.68, which Google positions for local codebase indexing, semantic code search and coding agent retrieval.
Weights are available on Hugging Face and Kaggle, and the model works with transformers, sentence-transformers, MLX, vLLM, llama.cpp, Ollama and LMStudio, and can run in the browser through transformers.js or WebGPU. For web developers, that means site search and related-content recommendations can run directly in the visitor's browser or device.