EmbeddingGemma 2: Open Multimodal Embedding Model
Google DeepMind released EmbeddingGemma 2, a 740 M‑parameter open model that embeds text, code, images, video and audio on device.
Original source published: October 6, 2026
EmbeddingGemma 2 is the latest open model from Google DeepMind, built on the Gemma 4 architecture and released under an Apache 2.0 license. With 740 million parameters, it supports native multimodal embeddings for text, code, images, video and audio, and is positioned as a lightweight option for on‑device inference.
The model achieves leading scores among sub‑1 B multimodal embedders on benchmarks such as MTEB Code and MAEB, and matches or exceeds larger models on text, vision and audio tasks. It is modular—text‑only workloads can run with as few as 270 M parameters, while optional vision (170 M) and audio (300 M) encoders enable full multimodal support. Matryoshka Representation Learning lets developers truncate vectors from 768 to 128 dimensions, offering up to 6× storage reduction. Quantized inference on a Google Pixel 11 Pro uses roughly 191 MB RAM for text‑only and 567 MB for the full model, and an 8K token window allows processing of up to 5.5 minutes of audio, 29 images or 58 video frames.
EmbeddingGemma 2 is intended for on‑device semantic search, routing and retrieval, preserving data privacy and reducing latency. It can be paired with Gemma 4 for on‑device retrieval‑augmented generation pipelines. Model weights are available on Hugging Face and Kaggle, with deployment options through Google AI Edge MediaPipe, LiteRT, transformers.js, WebGPU and various serving libraries.