computing• 3 min readOctober 6, 2026

Google DeepMind Releases EmbeddingGemma 2 for On-Device Multimodal AI

Google DeepMind launched EmbeddingGemma 2, a lightweight open model that handles text, images, video, and audio on local consumer devices. It builds on the Gemma 4 architecture and offers storage and memory efficiency for developers.

In short: Google DeepMind launched EmbeddingGemma 2, a lightweight open model that handles text, images, video, and audio on local consumer devices. It builds on the Gemma 4 architecture and offers storage and memory efficiency for developers.

Finding a specific video clip or audio recording on your personal device might soon get much easier and more private, thanks to new technology from Google DeepMind.

What happened, in plain words

Google DeepMind announced EmbeddingGemma 2, a new open and lightweight model designed to run on consumer hardware. Building on the Gemma 4 architecture and released under an Apache 2.0 license, the model combines text, code, images, video, and audio into a single shared space. It features 740 million parameters and allows developers to build local search tools and privacy-first retrieval augmented generation pipelines entirely on edge hardware.

Key points

  • Multimodal capabilities EmbeddingGemma 2 expands beyond text to unify code, images, video, and audio in a shared embedding space, allowing users to search through audio with text or find video clips from voice memos.
  • Size and efficiency The model has 740 million parameters and is modular by design, requiring as little as 270 million parameters for text-only workloads with optional encoders for vision and audio.
  • Storage and memory optimization Using Matryoshka Representation Learning, developers can truncate output vectors to reduce storage up to 6 times. With quantization on a Google Pixel 11 Pro, it requires about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.
  • Extended context window The model features an 8K token context window, which is four times larger than the previous version, letting it process up to 5.5 minutes of audio, 29 images, 58 video frames, or mixed combinations on local hardware.
  • Improved performance scores It delivers a 9.92-point improvement on code performance in MTEB Code, moving from 68.76 to 78.68, and matches strong multilingual text performance while outperforming some specialist models twice its size.

Terms explained

  • Embedding model — A type of artificial intelligence tool that translates different types of data, such as words or images, into numbers so computers can understand how pieces of information relate to each other. Example: Sorting a digital photo album by what is inside the pictures so you can easily find all photos of a beach.
  • Multimodal — Describing an artificial intelligence system that can process and understand multiple types of media at the same time, such as text, audio, and video. Example: An application that lets you speak a sentence out loud to find a matching scene in a movie.
  • Parameters — The internal adjustable settings inside an artificial intelligence model that help it learn patterns and make predictions. Example: Turning dials on a sound mixer to get the exact right balance of voice and music.
  • Quantization — A technique used to shrink the size of an artificial intelligence model so it takes up less memory and runs faster on regular devices. Example: Compressing a large digital photo file into a smaller format so it takes up less space on your phone.
  • Context window — The amount of information or data that an artificial intelligence model can look at and process all at once. Example: Reading a single chapter of a book at one time instead of scanning the whole book page by page.

Why it matters

EmbeddingGemma 2 helps developers build offline search tools and applications that run locally on consumer hardware. This keeps user data private on the device, works without an internet connection, and reduces latency for everyday tasks like searching through local codebases or personal media libraries.

What we still don't know

The source does not provide complete details on every possible real-world use case or long-term performance outcome across all consumer devices. Availability on the Gemini Enterprise Agent Platform Model Garden is coming soon rather than available immediately.


Based on reporting from Google DeepMind. This is an independent explainer, written in our own words with AI assistance; Google DeepMind has not reviewed or endorsed it. Read the original for the full details.

Google DeepMind Releases EmbeddingGemma 2 for On-Device Multimodal AI | Curious Tech Portal | Curious Tech Portal