Google Launches EmbeddingGemma 2 for Offline Multimodal Search
Google launched EmbeddingGemma 2 on October 6, releasing an open model that lets applications search text, code, images, video and audio on consumer devices. The 740-million-parameter system brings those formats into one shared numerical space, with weights available under the Apache 2.0 license.
The release expands Google’s earlier text-focused embedding model into a retrieval layer for mixed media. Its significance is practical: developers can build search that understands a spoken query or an image without first sending private files to a cloud service or converting every format into text.
EmbeddingGemma 2 Adds Search Across Media
An embedding represents content as numbers that software can compare for similarity. A search application embeds both the request and its stored files, then retrieves the closest matches. EmbeddingGemma 2 makes that comparison possible across formats, rather than maintaining separate search systems for each type of content.
Google’s launch announcement describes uses such as finding a video clip with a voice memo or searching audio recordings with a written query. The model can also supply retrieved material to a generative model through retrieval-augmented generation, commonly called RAG.
The model offers four encoder configurations:
- Text and code: 270 million parameters
- Text plus images and video frames: 440 million parameters
- Text plus audio: 570 million parameters
- All supported formats: 740 million parameters
The modular design means an application need not load an audio encoder when it only searches documents. Google reports approximately 191MB of active RAM for quantized text-only weights and 567MB for the full model on a Pixel 11 Pro. Those measurements describe a specific configuration, rather than a guarantee for every device.
Smaller Vectors Reduce Local Search Storage
The developer guide explains how all configurations produce compatible embeddings. A text query generated by the smallest setup can therefore search files indexed using the full model. Adding a modality later does not require recomputing previously generated embeddings.
Outputs normally contain 768 dimensions. Developers can shorten them to 512, 256 or 128 using Matryoshka Representation Learning, which trains the representation to retain useful information in its leading dimensions. Fewer dimensions mean less space for the index and less work during similarity comparisons.
Google recommends selecting the reduction according to the workload. Its guide positions 256 dimensions as a compromise for storage-constrained indexes, while 128 dimensions is better suited to large text collections or an initial shortlist followed by reranking. Multimodal retrieval loses more quality at the smallest size.
This makes the index an important deployment choice alongside the model itself. A compact set of weights does not eliminate the space required to represent a large media library. Developers still need to decide how much content to index, which dimensions to keep and how often to refresh stored representations.
Google AI Edge Adds Local Media Demonstrations
Google’s AI Edge announcement introduces Instant Media Search and Video Moments Finder in its Gallery app. The first retrieves local images and videos using natural language or example images; the second finds relevant moments inside selected recordings through cross-modal matching.
Google also launched the experimental AI Edge Foresight app for Mac. It combines EmbeddingGemma 2 with Gemma 4 for meeting notes, transcript recall and searches over private files. The company says its processing runs locally and can work offline, including with system audio and microphone input.
For production Android applications, Google plans ML Kit access in the coming weeks, including NPU acceleration where available. MediaPipe supplies higher-level embedding and retrieval tasks, while LiteRT offers a path for developers who need more control over processing and hardware acceleration.
These demonstrations show the intended product direction, but developers must test their own libraries and hardware. Search relevance, initial indexing time, storage consumption and battery use can determine whether an offline experience is useful even when the core model fits comfortably in memory.
Model Card Sets Benchmark and Safety Limits
The model card reports a code-retrieval MTEB score of 78.68, compared with 68.76 for its predecessor. Multilingual text scores change much less, from 61.15 to 61.36. These are Google’s full-precision evaluation results; quantized deployments need their own checks.
The system shares an 8,192-token input budget across modalities. Mixing text and media reduces the amount of each that fits. Google also cautions that performance is unequal across its supported languages and that recommended task prefixes influence text-retrieval quality.
EmbeddingGemma 2 has no post-training safety alignment or output moderation. Because it produces representations, Google places responsibility for retrieval filtering and fairness testing on application developers. Keeping computation local can limit data transmission, but access controls and safeguards around the resulting search remain necessary.
Background Reading