Feature Extraction
Transformers
Safetensors
sentence-transformers
multilingual
embedding_gemma2
embedding
multimodal-embedding
multimodal
vision
audio
video
image-feature-extraction
audio-feature-extraction
video-feature-extraction
sentence-similarity
Instructions to use google/embeddinggemma-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/embeddinggemma-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="google/embeddinggemma-2")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("google/embeddinggemma-2") model = AutoModel.from_pretrained("google/embeddinggemma-2", device_map="auto") - sentence-transformers
How to use google/embeddinggemma-2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("google/embeddinggemma-2") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Update README.md
#2
by hschechter - opened
README.md
CHANGED
|
@@ -36,6 +36,9 @@ library_name: transformers
|
|
| 36 |
|
| 37 |
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
|
| 38 |
|
|
|
|
|
|
|
|
|
|
| 39 |
EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:
|
| 40 |
|
| 41 |
* **Native multimodality:** Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
|
|
@@ -45,7 +48,7 @@ EmbeddingGemma 2 builds upon the architectural and capability advancements of Ge
|
|
| 45 |
* **Context length:** 8K token context window, capable of processing minutes of audio or video.
|
| 46 |
* **Task-steered representations:** Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).
|
| 47 |
|
| 48 |
-
###
|
| 49 |
|
| 50 |
| Parameters | Total | 740M |
|
| 51 |
| :---- | :---- | :---- |
|
|
@@ -317,10 +320,7 @@ In creating an open embedding model, we have carefully considered the following:
|
|
| 317 |
* Training data used for EmbeddingGemma 2 underwent safety filtering to mitigate the risk of these biases.
|
| 318 |
* **Misinformation and Misuse**
|
| 319 |
* Embedding representations can be misused to retrieve, classify, or otherwise organize embedded content in false, misleading or harmful ways.
|
| 320 |
-
* Guidelines are provided for responsible use with the model, see the [Responsible
|
| 321 |
-
* **Transparency and Accountability**
|
| 322 |
-
* This model card summarizes details on the model's architecture, capabilities, limitations, and evaluation processes.
|
| 323 |
-
* A responsibly developed open model offers the opportunity to share innovation by making Vision-Language Model (VLM) technology accessible to developers and researchers across the AI ecosystem.
|
| 324 |
|
| 325 |
### Risks Identified and Mitigations
|
| 326 |
|
|
|
|
| 36 |
|
| 37 |
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
|
| 38 |
|
| 39 |
+
|
| 40 |
+
## **Model Overview**
|
| 41 |
+
|
| 42 |
EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:
|
| 43 |
|
| 44 |
* **Native multimodality:** Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
|
|
|
|
| 48 |
* **Context length:** 8K token context window, capable of processing minutes of audio or video.
|
| 49 |
* **Task-steered representations:** Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).
|
| 50 |
|
| 51 |
+
### EmbeddingGemma 2
|
| 52 |
|
| 53 |
| Parameters | Total | 740M |
|
| 54 |
| :---- | :---- | :---- |
|
|
|
|
| 320 |
* Training data used for EmbeddingGemma 2 underwent safety filtering to mitigate the risk of these biases.
|
| 321 |
* **Misinformation and Misuse**
|
| 322 |
* Embedding representations can be misused to retrieve, classify, or otherwise organize embedded content in false, misleading or harmful ways.
|
| 323 |
+
* Guidelines are provided for responsible use with the model, see the [Responsible
|
|
|
|
|
|
|
|
|
|
| 324 |
|
| 325 |
### Risks Identified and Mitigations
|
| 326 |
|