Google has announced EmbeddingGemma 2, a 740-million-parameter multimodal embedding model designed to map text, code, images, video and audio into a shared embedding space. In a post on X, @googlegemma says the model is optimized for on-device use, with modular encoders, adjustable vector sizes, an 8K-token context window and an Apache 2.0 license. The announcement is available in Google Gemma’s X post.
The announcement matters for developers building local semantic search, code retrieval or retrieval-augmented generation (RAG) systems. EmbeddingGemma 2 is intended to turn different kinds of content into vectors—numerical representations that capture semantic relationships—so a system can compare a query with stored text, images, recordings or video segments without requiring an exact keyword match.
What EmbeddingGemma 2 can embed
An embedding model converts content into vectors that can be indexed and compared by similarity. In a conventional text-search system, a text query might retrieve documents whose wording or meaning is related to the query. A multimodal embedding model extends that idea across media types.
Google says EmbeddingGemma 2 places text, code, images, video and audio in a single, unified embedding space. That makes cross-modal retrieval possible in principle: a text query could be compared with visual or audio content, while an audio description could be used to locate a related video segment. Google gives the example of using a voice memo or text query to find a specific video clip, or searching audio recordings with a text query, as described in its detailed announcement.
The shared space does not mean that every application automatically understands every media library. Developers still need to divide content into searchable units, generate and store embeddings, retrieve candidate matches and decide how a downstream application should use those results. The model provides the common representation for those comparisons; the surrounding search or RAG pipeline determines how useful the final experience is.
A modular design for local deployment
EmbeddingGemma 2 has 740 million parameters in its full form, but Google describes a modular design that can reduce the components needed for a particular workload. The text-only configuration requires as little as 270 million parameters, with optional vision and audio encoders listed at 170 million and 300 million parameters respectively.
This allows a developer to choose a narrower deployment when an application only needs text or code retrieval, rather than loading the complete multimodal system. A media-search application that handles images, video and audio would need the relevant additional encoders, so its resource requirements would be higher than those of a text-only deployment.
Google also reports quantized memory figures for a Google Pixel 11 Pro: approximately 191MB of active RAM for text-only weights and approximately 567MB for the full multimodal model. These are Google’s reported figures for the stated device and configuration, not independent hardware testing or a guarantee for other phones, computers or deployment settings.
The model has an 8K-token context window, which Google describes as four times larger than the text-only EmbeddingGemma. The company says that capacity can cover up to 5.5 minutes of audio, 29 images, 58 video frames or interleaved combinations of those inputs on local hardware. The practical amount of usable content will still depend on how an application prepares and batches its inputs.
Why adjustable vector sizes matter
EmbeddingGemma 2 uses Matryoshka Representation Learning (MRL), a technique that allows the beginning portion of an embedding vector to remain useful when the full vector is shortened. Google says developers can truncate the model’s 768-dimensional output to 512, 256 or 128 dimensions.
Shorter vectors reduce the storage and memory needed by a local vector database. Google reports up to a sixfold reduction in storage and memory use, although the appropriate dimension depends on the application’s quality and resource requirements. A developer evaluating a deployment would need to test whether a smaller vector still preserves enough separation between relevant and irrelevant results for that dataset.
This gives local retrieval systems a deployment choice rather than one fixed vector size. Larger vectors may preserve more information, while smaller ones can make indexing and search more practical on constrained hardware.
What Google reports about retrieval quality
Google says EmbeddingGemma 2 delivers strong results for its size across text, code, vision and audio evaluations. The Google announcement cites the Massive Text Embedding Benchmark, or MTEB, as one of the evaluation suites; in its reported MTEB Code comparison, the model scores 78.68, compared with 68.76 for the original EmbeddingGemma—a stated improvement of 9.92 points. Google also highlights MIEB (Lite), a multimodal embedding evaluation, and MAEB, an audio embedding evaluation, alongside other tests.
The benchmark comparison supplied with the announcement shows EmbeddingGemma 2 alongside EmbeddingGemma, Jina v5 Omni-Nano, Qwen3-Embedding-.06B, SigLIP-So400M and larger-clip-general. The image marks some entries as unavailable, not supporting the relevant modality or not self-reported, so the columns are not a complete like-for-like result for every task. These remain Google’s reported evaluations rather than independent test results.

Image credit: @googlegemma on X
For developers, the code result is relevant to local codebase indexing, semantic code search and retrieval for coding agents. Those are proposed use cases supported by Google’s announcement, not evidence that the model will outperform every alternative on a particular repository or coding workflow.
Potential local search and RAG workflows
A basic semantic-search workflow could use EmbeddingGemma 2 to embed documents, source files, images or media segments, then store the resulting vectors in a local index. A user query would be embedded with the same system, and the application would retrieve the closest candidates for display or for a generative model to use as context.
The multimodal space enables several variations on that pattern:
A text query could search an image, video or audio collection.
A voice memo could help locate a related video segment.
Code embeddings could support repository indexing and semantic code search.
Local file retrieval could supply context to a separate generative model for an on-device RAG workflow.
Google points to Instant Media Search and Video Moments Finder in Google AI Edge Gallery, as well as the Foresight app, as examples of applications built around these ideas. It also describes using EmbeddingGemma 2 with Gemma 4 for local retrieval and contextual reasoning in its announcement. These are Google-described integrations and suggested workflows rather than independently tested products in this article.
Google describes the MediaPipe Decision Task API as a way to create real-time decision engines using multimodal context for classification, routing and predictive capabilities. The supplied announcement does not establish whether that API directly consumes EmbeddingGemma 2 embeddings, so it is best treated as a related deployment option rather than a confirmed model integration.
License, downloads and deployment options
Google says EmbeddingGemma 2 is released under the commercially permissive Apache 2.0 license. The linked announcement lists model weights on Hugging Face and Kaggle, with availability in Gemini Enterprise Agent Platform Model Garden described as coming soon at the time of publication.
For on-device development, Google lists MediaPipe for turnkey embedding, retrieval and decision tasks, and LiteRT for custom model integration. The announcement also names browser deployment through transformers.js or WebGPU and lists tools including Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. It identifies Qdrant as one option for storing embedding vectors and Unsloth as a source of fine-tuning guidance.
The list describes the deployment and tooling ecosystem Google says it worked with; it does not establish that every tool has identical support, compatibility or performance for every model configuration. Developers should confirm the relevant model format, encoder support, quantization path and runtime requirements before choosing a deployment.
EmbeddingGemma 2 therefore combines a shared multimodal representation with deployment controls aimed at local hardware. Its adjustable encoders and vector sizes may help developers fit retrieval systems to different resource budgets, while its reported benchmark results provide Google’s case for using it across code, text and media. Independent tests would still be needed to determine the best configuration for a specific dataset, device and retrieval quality target.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment