Gemini Embedding 2 is Google's embedding generation model that's ideal for complex retrieval and analytics tasks.
Gemini Embedding 2 accepts multimodal inputs to generate 3072-dimensional vectors. It accepts images, text, documents, audio, and video inputs and semantically maps the generated vectors into a unified semantic space. This lets you perform tasks, such as searching for an image based on a text description.
Gemini Embedding 2 introduces several features to optimize embedding quality and flexibility:
Custom task instructions: By specifying task instructions (for example,
task:code retrievalortask:search result) optimize the embeddings for the intended relationships and retrieve more accurate results for the specific goal.Adjustable result size: The model generates a 3072-dimensional float vector, by default. However, you can retrieve a smaller dimensional output by specifying the
output_dimensionalityparameter.Document OCR: Read OCR from document inputs.
Audio track extraction: Extract audio tracks from video inputs and interleave them with video frames.
For more information on how to use Gemini Embedding 2, see Get multimodal embeddings.
| Model ID | gemini-embedding-2 |
|
|---|---|---|
| Modalities |
|
|
| Token limits | Maximum input tokens | 8,192 |
| Maximum output tokens | N/A | |
| Output dimensions | Up to 3,072 (with MRL support) | |
| Maximum sequence length | 8,192 tokens | |
| Consumption options |
|
|
| Technical specifications | Text |
|
| Image |
|
|
| Video |
|
|
| Audio |
|
|
| Supported regions |
|
|
| Knowledge cutoff date | November 2025 | |
| Versions |
|
|