Home / Companies / Google Cloud / Blog / Post Details
Content Deep Dive

EmbeddingGemma 2: The Developer Guide

Blog post from Google Cloud

Post Details
Company
Date Published
Author
Maarten Grootendorst, and Ian Ballantyne
Word Count
1,385
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

EmbeddingGemma 2 is a compact Apache 2.0–licensed multimodal embedding model designed for efficient search and retrieval-augmented generation across text, code, images, video, and audio. Built on Gemma 4, it maps all supported inputs into a shared 768-dimensional vector space and uses modular encoders, allowing deployments ranging from a 270-million-parameter text-and-code configuration to a 740-million-parameter full multimodal model. It improves code and technical retrieval over EmbeddingGemma 1, supports cross-modal similarity search and interleaved media inputs, and can be used through sentence-transformers and other common inference tools. Its Matryoshka Representation Learning capability permits embeddings to be truncated to 512, 256, or 128 dimensions to reduce vector-storage needs, with 256 dimensions retaining most quality for text and code and about 95% for visual and audio retrieval. The model supports an 8,192-token shared context window, can reuse existing embeddings when additional modality encoders are enabled, and reportedly scores 14% higher than its predecessor on the MTEB Code benchmark while preserving multilingual text retrieval performance.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 12 No monthly metrics for this publish month.
RAG 3 No monthly metrics for this publish month.
AI Model Fine-tuning 1 No monthly metrics for this publish month.
MLX 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.