Google DeepMind Releases 740M-Parameter Embedding Model Under Apache 2.0

Google DeepMind launched EmbeddingGemma 2 on October 6, 2026, an open embedding model with 740 million parameters that extends its predecessor from text-only to a shared embedding space covering code, images, video, and audio.

Five different materials — paper, metal, color, glass and wood — placed on a coordinate grid, illustrating a shared embedding model for text, code, image, video and audio.
Illustration
Gift article

Google DeepMind Releases 740M-Parameter Embedding Model Under Apache 2.0

Google DeepMind launched EmbeddingGemma 2 on October 6, 2026, an open embedding model with 740 million parameters that extends its predecessor from text-only to a shared embedding space covering code, images, video, and audio.

What Changed

The original EmbeddingGemma was a text model. Version 2, built on the Gemma 4 architecture, is intended by Google to bring all modalities together in a single embedding space. In practice, that means text search no longer necessarily has to run as a separate, text-based pipeline: an embedding of an audio clip, a video frame, or a piece of code can, in principle, be searched against the same vector space as text.

The announcement came in a blog post by research engineers Sahil Dua and Henrique Schechter Vera, who write (quote translated from English): "Today we're launching EmbeddingGemma 2, which goes beyond text to bring code, images, video and audio together in a shared embedding space" (Google DeepMind).

The model's weights were made available immediately on Hugging Face and Kaggle under the Apache 2.0 license, which permits commercial use. Alongside the model, Google released a macOS app called AI Edge Foresight, a demo showing how the model can run fully locally on a laptop.

Modular Architecture and Context

The model has 740 million parameters in total but is built modularly. For text-only workloads, only a 270-million-parameter component is required, and full multimodal support adds an optional 170-million-parameter vision encoder and a 300-million-parameter audio encoder. That is how Google can state that the full multimodal model still fits on-device.

The context window is 8K tokens, four times the size of EmbeddingGemma 1. According to Google, it accommodates up to 5.5 minutes of audio, 29 images, 58 video frames, or combined combinations of these, processed directly on local hardware.

Another practical measure is Matryoshka Representation Learning (MRL). The output vectors have 768 dimensions, but developers can dynamically truncate them to 512, 256, or 128 dimensions. Google claims this yields up to 6x storage savings for local vector databases and memory use — relevant because storage and memory are the hardest constraints for search and retrieval systems on mobile.

The Numbers — Google's Own Claims

All performance figures below come from Google itself and have not been independently verified as of the sources' publication dates:

  • RAM requirements with quantization: On a Google Pixel 11 Pro, Google says the model requires as little as roughly 191 MB of active RAM for text-only weights and roughly 567 MB for the full multimodal model.
  • Code quality: Google claims a 9.92-point improvement on the MTEB Code benchmark, from 68.76 to 78.68, while multilingual text performance reportedly matches the predecessor.
  • Class leadership: Google presents the model as best in class among multimodal embedding models under one billion parameters, measured by quality per parameter.

These claims should be read as the company's own measurements, not established results. The MTEB benchmark is a well-known framework, but the numbers are not confirmed by third parties in the available source material.

The AI Edge Foresight Demo App

To showcase the model, Google simultaneously launched AI Edge Foresight, a free macOS app for Apple Silicon. According to Android Authority's reporting — including a hands-on by journalist Akshay Gangwar — the app turns short keyword notes into fully formatted notes during meetings, and can answer questions about the meeting's content while it is ongoing. Everything is said to happen fully locally on the machine, with no data sent to the cloud (Android Authority).

It is worth noting that the app's details come from Android Authority's coverage, not from primary Google documentation in the available source material. According to Android Authority, Google itself describes the model as designed for "privacy-first" applications and as an "ultra-low-latency on-device decision engine."

Why It May Matter

The core of the launch is not the benchmark numbers but the distribution model: the Apache 2.0 license permits commercial use, and the weights are publicly available. For developers building RAG (retrieval-augmented generation) solutions and vector-based search, it means the entire retrieval infrastructure — both embedding and vector store — can in principle run offline on a phone or laptop. Three things combine to make this possible: the modular architecture that keeps the RAM requirement down, MRL truncation that shrinks the vector database, and the open license that removes licensing barriers for commercial products.

The interest is not theoretical. According to Cryptobriefing, the predecessor EmbeddingGemma passed 20 million downloads, suggesting an existing developer community around this model family (Cryptobriefing).

Open Questions

The key uncertainties are worth keeping clearly in view:

  • Unverified performance. The MTEB Code figures, the best-in-class claim, and the RAM measurements are Google's own. Until independent benchmarking exists, they are promises of quality rather than documented results.
  • Real-world quality. A shared multimodal embedding space is attractive in theory, but it remains to be seen how well mixed searches — for example, text queries against audio or video — work in real applications. The source material contains no such tests.
  • Demo versus platform. AI Edge Foresight is a showcase app built on secondary reporting. How much of that experience can be reused in developers' own products is not yet known.
  • Small framing differences. Cryptobriefing describes the 270M component as dedicated to "text and code," while Google itself frames it as the text-only configuration. It is a minor difference that does not contradict the modular architecture, but it illustrates that details should still be checked against Google's own documentation at implementation time.

With 20 million downloads of its predecessor, a commercially permissive license, and concrete numbers for on-device operation, Google has laid favorable groundwork for rapid adoption. Whether the model holds up in practice will have to be shown by independent measurement.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.