Google DeepMind has launched EmbeddingGemma 2, an open, multimodal embedding model designed to process text, code, images, video, and audio on consumer hardware. The model, built on the Gemma 4 architecture, offers a unified embedding space and can run entirely on devices without an internet connection, enhancing privacy and reducing latency.
Information was available with The Chenab Times that EmbeddingGemma 2 features 740 million parameters and is released under a commercially permissive Apache 2.0 license. This new model expands upon the capabilities of its predecessor, EmbeddingGemma, which focused solely on text embeddings and has seen over 20 million downloads since its release in 2025. The enhanced model aims to enable developers to build more sophisticated on-device applications, including privacy-first retrieval-augmented generation (RAG) pipelines.
EmbeddingGemma 2 is engineered with a modular architecture. It includes a 270-million-parameter text model, with optional vision and audio encoders of 170 million and 300 million parameters, respectively. This modularity allows developers to load only the necessary components, scaling the model’s footprint from 270 million parameters for text and code to 740 million for full multimodal processing. All configurations project into a single, unified 768-dimensional vector space.
The model boasts an 8K token context window, which is four times larger than that of EmbeddingGemma 1. This extended context allows it to process significant amounts of data, such as up to 5.5 minutes of audio, 29 images, or 58 video frames. With quantization on devices like a Google Pixel 11 Pro, EmbeddingGemma 2 can operate with as little as approximately 191 megabytes of active RAM for text-only weights and around 567 megabytes for the full multimodal model.
Google DeepMind highlights that EmbeddingGemma 2 achieves top-tier performance for its size across various benchmarks, including code, vision, and audio tasks. It demonstrates a significant improvement in code performance, with a 9.92-point increase on the MTEB Code benchmark, moving from 68.76 to 78.68. The model also matches or surpasses some specialist models that are more than twice its size in terms of quality per parameter.
The release of EmbeddingGemma 2 is expected to foster innovation in on-device AI applications. Developers can leverage its capabilities for tasks such as finding specific video clips from voice memos, searching extensive audio recordings using text queries, and indexing local codebases for semantic search. The model’s availability on platforms like Hugging Face and Kaggle, along with its permissive license, facilitates widespread adoption and experimentation by the developer community.
❤️ Support Independent Journalism
Your contribution keeps our reporting free, fearless, and accessible to everyone.
Or make a one-time donation
Secure via Razorpay • 12 monthly payments • Cancel anytime before next cycle


(We don't allow anyone to copy content. For Copyright or Use of Content related questions, visit here.)

The Chenab Times News Desk





