Ground Truth.
AI, checked against the source.

News · 2026-10-07

Google releases EmbeddingGemma 2 for local search across words, pictures, and sound

Google released EmbeddingGemma 2 on October 6 as an open-weight model for searching across text, code, images, audio, and sampled video in a shared representation space. Google reports that its quantized full configuration uses about 567 MB of active memory on a Pixel 11 Pro. The release makes private local media search more practical for developers, while leaving ingestion, indexing, and the user interface to applications.

Key facts

Finding a photograph with a sentence or locating the right moment in a voice memo requires a common way to compare unlike things. A filename search cannot easily match a spoken recollection to an uncaptioned picture. EmbeddingGemma 2 supplies numerical representations that allow an application to compare those inputs by meaning rather than by identical strings.

The model card describes a modular design: a 270-million-parameter text component, an optional 170-million-parameter vision encoder, and an optional 300-million-parameter audio encoder. Together they make the 740-million-parameter configuration. An application can load the encoders it needs rather than treat every search task as requiring the full set.

An embedding is a list of numbers that places an item in a learned space. Imagine a map where a photo of a bicycle, the words describing a bicycle, and a related audio clip can receive nearby coordinates. A query also receives coordinates. Search then compares locations and returns nearby indexed items. The analogy explains the operation without implying that distance proves factual identity or that the model understands every detail of the source.

The existing embeddings lesson introduces that foundation. Version 2 extends the earlier text model to cross-media retrieval. Its practical novelty is comparable representations across modalities, rather than a chatbot producing prose about every file. Google’s developer guide gives examples of comparing a text query with content represented using the appropriate media encoders.

Google’s edge guide calls the use case “multimodal semantic search.” The AI Edge deployment post describes local demos in AI Edge Gallery, including Instant Media Search and Video Moments Finder. The latter indexes visual frames and audio chunks, then returns matching timestamps. That is a working demonstration route, not a promise that the model automatically becomes a phone-wide search service.

Memory figures need careful labeling. The reported 191 MB and 567 MB are active memory measurements after quantization on a Pixel 11 Pro. They are neither download sizes nor general requirements for graphics-card memory. The dossier contains no verified weight-file total, and the synthesis’s single attempted file-listing check was inaccessible. The minimum or recommended graphics-memory requirement is unstated in the reviewed sources. It would be misleading to calculate it from the parameter count or reuse the phone’s memory measurement as a universal requirement.

The card’s general inference advice also differs from the quantized phone demonstration: it recommends bfloat16 or float32 and warns that float16 can create invalid or degraded embeddings. A developer needs the right format and deployment path, not simply any lower-precision load. Quantization explains why reduced numerical precision can change behavior as well as resource use.

Inputs share an 8,192-token budget. The card gives approximate single-modality limits of 29 images, 58 video frames, or about five and a half minutes of audio at default settings. Mixed inputs compete for the same space. Video defaults to one frame per second; this is sampled retrieval, not continuous full-frame analysis. Audio should be mono at 16 kilohertz. The model does not itself return transcripts, extracted document text, summaries, or an end-user search interface.

Google also allows shorter stored vectors through Matryoshka representation learning. It reports little quality loss down to 256 dimensions, while 128 dimensions substantially hurts multimodal quality. Reducing vector length shrinks the index, not the neural-network weights. After truncation the vectors must be normalized again. That distinction becomes especially valuable when an application stores representations for a large personal library.

The maker reports a 9.92-point improvement on its code-retrieval comparison, with multilingual text results broadly similar to version 1. Those are Google’s benchmarks, not independent replication. The Hacker News discussion shows enthusiasm for local search and the Apache license, alongside accuracy and index-maintenance questions; one user reports an unsuccessful example. Early reactions establish interest rather than production reliability.

Google labels the checkpoint Apache 2.0, while the card separately requires adherence to its prohibited-use policy. The published model page is the artifact route. The honest deployment question is whether retrieval quality, file updates, and permission handling work for the intended collection. Local execution can keep media from being uploaded, but privacy and usability still depend on the surrounding application.


Primary source, verified: read the paper →

Key questions

Is 567 MB the EmbeddingGemma 2 download size?

No: Google reports about 567 MB of active RAM for its quantized full configuration on a Pixel 11 Pro. The dossier does not establish the checkpoint’s disk size or a general GPU-memory requirement.

Can EmbeddingGemma 2 answer questions about my photo library by itself?

It returns embeddings rather than generated answers or a finished index. An application must ingest files, maintain a search index, and present retrieved results.

Does its video search examine every frame?

No: the model card’s default is one sampled video frame per second. Sampling is configurable, and text and media share one input budget.
Cite this

APA

Ground Truth. (2026, October 7). Google releases EmbeddingGemma 2 for local search across words, pictures, and sound. Ground Truth. https://groundtruth.day/news/embeddinggemma-2-multimodal-on-device-search.html

BibTeX

@misc{groundtruth:embeddinggemma-2-multimodal-on-device-search,
  title  = {Google releases EmbeddingGemma 2 for local search across words, pictures, and sound},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/news/embeddinggemma-2-multimodal-on-device-search.html}
}

Topics: embeddings · on-device-ai · multimodal · google · open-weights

Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.