Google EmbeddingGemma 2 Brings Multimodal Search On-Device
Google released EmbeddingGemma 2 on 6 October 2026, an open 740 million-parameter model that embeds text, code, images, video and audio in one shared space for on-device search.
PromptCrates Editorial
Staff Writer

Google released EmbeddingGemma 2 on Tuesday, 6 October 2026, an open embedding model with 740 million parameters that places text, code, images, video and audio in one shared vector space. Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera wrote in the announcement on Google's blog that the model is built on the Gemma 4 architecture and released under the commercially permissive Apache 2.0 license. Its predecessor, the text-only EmbeddingGemma, has passed 20 million downloads, according to the post.
Embedding models turn content into lists of numbers so that software can measure how similar two items are. They sit underneath semantic search, recommendation and retrieval augmented generation, or RAG, where an app fetches relevant material before a language model answers. Google is aiming the new version at on-device use, saying that generating embeddings locally helps protect data privacy and reduces pipeline latency.
How the modular encoders fit together
The release is built from parts that can be loaded separately. According to Google's developer guide, the core is an adapted Gemma 4 decoder with an 8,192-token context window, joined by a vision encoder for images, visual documents such as PDFs, slides and charts, and video frames, plus a dedicated encoder that ingests raw speech and sound.
Developers choose which pieces to load. The guide lists four setups: 270 million parameters for text and code, 440 million for text plus vision, 570 million for text plus audio, and the full 740 million for every modality. All four load from the same checkpoint and project into the same 768-dimensional space, so a query embedded with the small text-only setup can be matched directly against documents embedded with the full model. The guide says a team that starts with a text-only index and later adds image or audio embeddings does not need to recompute the embeddings it already has.
Google's blog post says the 8K context is four times larger than the first EmbeddingGemma's and lets the model take in up to 5.5 minutes of audio, 29 images or 58 video frames, or mixtures of them, in a single input. The developer guide shows interleaved inputs, such as a product listing that combines a text description, a photo and a short video clip into one embedding.
Benchmarks and memory figures Google reports
The headline quality number is for code. Google says EmbeddingGemma 2 improves on its predecessor by 9.92 points on MTEB Code, the code portion of the Massive Text Embedding Benchmark, moving from 68.76 to 78.68, while matching the first model's multilingual text performance. The company says the model achieves leading scores among multimodal embedders under 1 billion parameters on MTEB Code and on MAEB, the Massive Audio Embedding Benchmark, and that it outperforms some specialist models more than twice its size. These are Google's own results; the post points readers to the model card for full evaluation metrics.
For on-device use, Google gives memory figures measured on a Google Pixel 11 Pro with quantization: about 191MB of active RAM for the text-only weights and about 567MB for the full multimodal model.
Storage is the other constraint the release targets. EmbeddingGemma 2 uses Matryoshka Representation Learning, which lets developers truncate each vector from 768 dimensions to 512, 256 or 128. The developer guide gives a worked example: a million 768-dimensional vectors in bfloat16 precision take roughly 1.5GB, while the same million at 128 dimensions need about 250MB, a sixfold reduction. At 256 dimensions, the guide says, the model keeps most of its full quality on text and code and about 95 percent on image, video and speech retrieval.
Where developers can run the model
Google says the weights are available on Hugging Face and Kaggle, with availability in the Gemini Enterprise Agent Platform Model Garden listed as coming soon. Versions optimized for on-device use sit in the LiteRT Community on Hugging Face. For apps, the post points to Google AI Edge MediaPipe for embedding, retrieval and decision tasks, LiteRT for custom integration, and transformers.js or WebGPU for the browser.
The post lists support across transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio, vector storage through Qdrant, and fine-tuning guidance from Unsloth. Google also published demos: Instant Media Search and Video Moments Finder in the Google AI Edge Gallery, which find media from text, image or audio queries, and the Google AI Edge Foresight app, which pairs EmbeddingGemma 2 for local file retrieval with Gemma 4 for reasoning. Because the two models share a text tokenizer and audio encoder, Google says running them together lowers the combined memory footprint.
In one example from the developer guide, Google embedded the Hugging Face transformers codebase using the 270 million parameter text-only setup, then let an agent built on Gemma 4 26B A4B with the Pi agent harness search it by semantic similarity.
What the announcement leaves open
The benchmark claims in the announcement are Google's own measurements, and the company directs readers to the EmbeddingGemma 2 model card for the full evaluation tables. The hosted route is also not finished: Google lists Model Garden availability as coming soon rather than live at launch.
The release extends Google's recent work on multimodal and agent tooling that PromptCrates has covered, including Gemini's agentic video understanding, which cut tokens by up to 88 percent, and WikiSkill, which gives agents persistent memory skills. EmbeddingGemma 2 differs from those projects in where it runs: Google says the whole pipeline, from embedding to search, can work entirely offline on local hardware.


