NewsGoogle2 min read
EmbeddingGemma 2: Search by Meaning in Text, Images and Audio
Google releases EmbeddingGemma 2, a small open embedding model that understands text, code, images, video and audio, and runs offline on a phone.
Key points
- It turns text, code, images, video and audio into vectors that can be compared with each other.
- At 740 million parameters it fits on a phone: about 570 MB of memory for the full version.
- It is open (Apache 2.0 licence) and downloadable from Hugging Face, Kaggle or Ollama.
- It reads up to 8,000 tokens at once, four times more than the first version.
On 6 October Google released EmbeddingGemma 2, the second version of its open embedding model. It is not a chatbot: it is a building block many AI apps use under the hood, and this version makes it more useful and easier to run on any device.
What an embedding model does
An embedding is a list of numbers that captures the meaning of a piece of content. Two sentences that say the same thing in different words end up with similar numbers. That is how “search by meaning” works, and how RAG lets a chatbot answer from your own documents: the documents are turned into embeddings, stored in a vector database, and the closest ones to the question are retrieved.
What version 2 brings
- Multimodal: text, code, images, video and audio share the same space. You can find a photo by typing a sentence, or an audio clip from a description.
- Small: 740 million parameters in total, or 270 million if you only need text. Google says it uses about 191 MB of memory on a phone (text only) or 567 MB (full), so it works offline, without sending data anywhere.
- More context: up to 8,000 tokens per input, four times the first version. That is about 5.5 minutes of audio or 29 images.
- Better at code: according to Google, almost 10 points higher on the MTEB Code benchmark.
Where to use it
It is open, under the Apache 2.0 licence, which allows commercial use. You can download it from Hugging Face and Kaggle, and it works with the usual tools for running models locally, such as Ollama, llama.cpp or MLX, and even in the browser with transformers.js.
What it means for you
- If you are building RAG or an internal search engine, it is a free and private option: your documents never leave your computer.
- If you work with photos, video or audio, you can now search them by content with a single model instead of one per format.
- If you only use chatbots, you will not see it directly, but it is the kind of component that makes searching your files work better.


