Google's New Open Embedding Model Puts Multimodal Search on a Phone. Before You Index Scanned PDFs With It, Check the Image Budget

On October 6, Google DeepMind released EmbeddingGemma 2, an open embedding model that maps text, code, images, video and audio into one shared 768-dimensional space. It has 740 million parameters, it is built on Gemma 4, and it ships under Apache 2.0. The model card says it is designed for consumer hardware such as phones and laptops, and Google's on-device card lists timings on an iPhone.
I read the launch post, the model card and the on-device card that evening with one question in mind. I build a LegalTech that handles patent documents, and I ship iOS apps on the side. In both worlds, the first question about an embedding model is rarely "how good is it?". It is "where does the document go when I index it?".
What Google actually shipped
The model is modular. The text backbone is 270M parameters (130M of transformer, 140M of embedding table). A 170M vision encoder and a 300M audio encoder are optional. Text plus vision comes to 440M; everything, 740M. All configurations share one vector space, so a query embedded with the text-only setup can be matched against pages embedded with the full model.
The context window is 8,192 tokens, four times the first EmbeddingGemma. An image costs 280 tokens by default, so about 29 images fit in one input. The card claims more than 100 languages. The output is trained with Matryoshka Representation Learning: you can keep only the first 512, 256 or 128 dimensions, for up to six times less storage.
The license changed too. The first EmbeddingGemma sits on Hugging Face under the "gemma" license, behind a gate where you agree to Google's usage terms. Version 2 is tagged apache-2.0, with no gate. The model card still says deployments must follow the Gemma Prohibited Use Policy, which a legal team will want to read next to the license.
The numbers, and what they are compared to
The model card compares EmbeddingGemma 2 with exactly one model: its predecessor. Multilingual text is flat (61.36 against 61.15 on MTEB multilingual v2). Code jumps from 68.76 to 78.68 on MTEB Code, the 9.92 points of the headline. Everything else is new, so it has nothing to compare with: 64.64 on MIEB lite for images, 50.67 on MMEB v2 video, 69.54 on MSEB audio retrieval.
The launch post adds three charts that place it by size against open models: Qwen3-Embedding 0.6B and 8B for code, jina-embeddings-v5-omni and LCO-Embedding-Omni for images and audio, among others. The top of each chart belongs to a much larger model, Qwen3-Embedding-8B for code and LCO-Embedding-Omni (3B, then 7B) for images and audio. The claim is quality per parameter, and the charts are drawn to show exactly that. Google's own hosted model, Gemini Embedding 2, appears on none of them.
The number I care about is another one: MMEB v2 VisDoc, visual document retrieval, 67.84 NDCG@5. It is the closest thing on the card to "find the right scanned page". It has nothing next to it.
Why a local multimodal embedder matters for documents
Most of the documents I deal with are not clean text. Patent files mix claims with drawings. Prior art arrives as scanned PDFs. A figure with reference numerals carries meaning that OCR flattens or drops. The usual pipeline chains OCR, sometimes a captioning model, then a text embedder, and every link is a place where something breaks or leaves your machine. Google's edge team describes the same problem for media: one model "reduces the latency and memory overhead of chaining separate image captioning, speech-to-text, and text-embedding models".
An embedding is also not anonymous at creation time. To compute it, a hosted service has to receive the whole document. When the question is "who sees this invention before it is filed?", "a large cloud vendor, under contract" is an answer. "Nobody, it runs on our machine" is a better one. A 740M model under Apache 2.0 can run on your own server, in the region you chose, behind sentence-transformers, vLLM, llama.cpp or Ollama, all of which Google lists as supported.
The cost argument is weaker than you think
I checked Google's hosted equivalent the same day. Gemini Embedding 2 takes text, images, video, audio and PDFs. On the paid tier, an image costs $0.00012, or $0.00006 in batch. A hundred thousand scanned pages comes to about $12, or $6 in batch. Nobody should change architecture to save $12.
Two other lines on that pricing page matter more. The free tier is marked "Used to improve our products: Yes". The paid tier is marked "No". My side projects usually start on free tiers, and that is the line to read before indexing anything a user handed you.
The real cost is the one Simon Willison raised on Hacker News the day of the launch. Embeddings are computed by the thousand or the million and stored. If a vendor retires its model, you pay to re-embed everything. His nuance is worth keeping: he would rather pay a provider to host it, knowing the open weights are there if the provider stops. Open weights do not force you to self-host. They give you an exit.
What the phone version gives up
For iOS, Google's LiteRT card publishes real timings. On an iPhone 18 Pro GPU, a text embedding takes 11.6 ms and text plus an image 69.8 ms, with 85 MB and 196 MB of memory. The text-and-vision bundle is a 388 MB download, which your App Store page will feel.
Now read the footnotes. Those timings use a 128-token text input and a 70-token image. The LiteRT bundles accept image budgets of 70 or 140 tokens only. The model card's default is 280 and goes up to 1,120, and it says a larger budget improves "fine-grained visual understanding". A holiday photo may survive 70 tokens. A dense page of claims with small reference numerals may not. The benchmarks were run on the full-precision checkpoint, and neither card gives a document retrieval score for the phone configuration.
So my split would be: documents embedded server-side at the default budget or higher, the phone handling queries and light media search. Both land in the same vector space, which is what makes the split possible at all.
What I would do on Monday
Build a small evaluation set from your own documents before choosing anything: fifty real queries and the pages that should come back. Run it at 768, 512 and 256 dimensions, and at two image budgets. Google's developer guide recommends 768 or 512 for visual document retrieval, and says 128 drops image and speech retrieval to around 75 percent of full quality.
Store the configuration with every vector: model, dimension, image budget, task prefix. Vectors from different settings are not comparable, and one day you will need to know which ones to redo.
Mind the three traps the model card spells out. Never run it in float16: its activations exceed the format's range, and you get NaN or silently degraded embeddings instead of an error. Re-normalize after truncating, or ranking quality degrades "silently". And use the task prefixes: SearchQuery for queries, "title: {title} | text: {content}" for documents.
The API itself is small. This is the developer guide's pattern with sentence-transformers 6.1 or later, loading text and vision only:
from sentence_transformers import SentenceTransformer
# Text, images, and video (440M parameters)
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"audio_config": None},
)
page_emb = model.encode({"image": "scan_page_017.png"})
query_emb = model.encode("hinged lid with a magnetic latch", prompt_name="SearchQuery")
print(model.similarity(query_emb, page_emb))
Last rule, the one that outlives this model: keep the source pages, not just the vectors. They are your only real protection against the next model, open or not.
Sources
- Google, "EmbeddingGemma 2: an open, lightweight multimodal embedding model" (October 6, 2026)
- Google AI for Developers, "EmbeddingGemma 2 model card" (last updated October 6, 2026)
- Google Developers Blog, "EmbeddingGemma 2: The Developer Guide" (October 6, 2026)
- Google Developers Blog, "Bring multimodal semantic search to the edge with EmbeddingGemma 2" (October 6, 2026)
- Hugging Face, "litert-community/embeddinggemma-2-740m-litert-lm" (read October 7, 2026)
- Hugging Face, "google/embeddinggemma-2" (read October 7, 2026)
- Hugging Face, "google/embeddinggemma-300m" (read October 7, 2026)
- Google AI for Developers, "Gemini Developer API pricing" (last updated October 6, 2026)
- Google AI for Developers, "Embeddings" (last updated September 17, 2026)
- Simon Willison, "Comment: EmbeddingGemma 2" (October 6, 2026)
