Google DeepMind 推出 EmbeddingGemma 2:740M 參數開放多模態嵌入模型,基於 Gemma 4 打造
Google DeepMind has released EmbeddingGemma 2, an open model that embeds text, code, images, video and audio into one 768-dimensional space.It has 740M parameters, an 8K token context window and an Apache 2.0 license.It targets on-device search, classification and privacy-first RAG.
This article analyzes, compares and showcase how EmbeddingGemma 2 fits in the space.Deployable today?Yes.Weights are live on Hugging Face and Kaggle, with Ollama, llama.cpp GGUF and LiteRT builds available now.
What an Embedding Model Does An embedding model converts content into a vector of numbers that captures meaning.Similar items land close together, so they are easy to search and compare.In a RAG pipeline, these vectors let an LLM retrieve fresh information it was not trained on.
Generating embeddings locally keeps data on the device, cuts latency and works offline.One Vector Space for Every Modality EmbeddingGemma 2 is built on the Gemma 4 architecture.A text query can retrieve a photo.A voice memo can retrieve a video clip.
Interleaved inputs, like a product listing with text, images and a demo video, produce a single embedding.The design is modular.
It has three parts: Text and code backbone: 270M parameters (130M transformer plus 140M embedder) Vision encoder: 170M parameters, optional Audio encoder: 300M parameters, optional Developers load only what they need: 270M for text, 440M for text and vision, 570M for text and audio, or 740M for everything.
All setups share one vector space.A query embedded with the text-only setup can match documents embedded by the full model.The context window is 8,192 tokens, 4x larger than version 1.That fits about 29 images, 58 video frames or 5.5 minutes of audio.
Benchmarks Google research team reports leading scores among sub-1B multimodal embedders on MTEB Code and MAEB.Full-precision results at 768 dimensions: BenchmarkEmbeddingGemma 2EmbeddingGemma 1MTEB multilingual v261.3661.15MTEB Code v178.6868.76MIEB lite (image)64.64n/aMMEB v2 overall59.
01n/aMSEB retrieval (sound)69.54n/aMAEB (audio)49.39n/a Source: EmbeddingGemma 2 model card Code retrieval gains 9.92 points, roughly 14%.Multilingual text quality holds steady.Bigger models still lead some boards.Qwen3-VL-Embedding-2B reports 73.2 on its own MMEB-V2 run, with about 2.
7x the parameters and no audio support.Built for Phones and Laptops With quantization on a Pixel 11 Pro, active RAM is about 191MB for text-only weights.The full multimodal model needs about 567MB.Quantization-aware training compresses weights to INT4 and INT8.The Google AI Edge team measured 37.
3 ms per image on a MacBook M5 Pro GPU, using a 70-token vision budget.Matryoshka Representation Learning (MRL) lets developers truncate vectors to 512, 256 or 128 dimensions.Moving from 768 to 128 dimensions cuts storage up to 6x.At 256 dimensions, MTEB multilingual only slips from 61.36 to 60.41.
At 128 dimensions, MMEB drops to 45.65, so Google recommends 128d mainly for text-only workloads.Interactive Explainer (function(){var f=document.getElementById('mtp-eg2-frame');window.addEventListener('message',function(e){if(f&&e.source===f.contentWindow&&e.data&&e.data.eg2h){f.style.height=e.
data.eg2h+'px';}});})(); EmbeddingGemma 2 vs.
Closest Competitors FeatureEmbeddingGemma 2EmbeddingGemma 1Qwen3-VL-Embedding-2BLCO-Embedding-Omni-3BGemini Embedding 2DeveloperGoogle DeepMindGoogle DeepMindAlibaba QwenLCO-Embedding (research)GoogleParameters740M (270M text-only)308M2B3B backbone (5B listed on HF)Not disclosedText / codeYesYesYesYesYesImagesYesNoYesYesYesVideoYesNoYesYesYesAudioYesNoNoYesYesOutput dims (MRL)768 (512, 256, 128)768 (down to 128)Up to 2048 (64 to 2048)Not stated3072 (128 to 3072)Context8,192 tokens2K tokens32K tokensNot stated8,192 tokensLanguages100+100+30+Not stated100+License / accessApache 2.
0, open weightsOpen weights (Gemma terms)Apache 2.0, open weightsApache 2.0, open weightsPaid API onlyPublished on-device RAM~191MB text, ~567MB fullUnder 200MBNot publishedNot publishedCloud onlySourceModel cardDocsHF cardHF cardAPI docs How to Run It It runs on sentence-transformers v6.1.
0+, Transformers, vLLM, SGLang, MLX, llama.cpp, Ollama, LM Studio, LiteRT and MediaPipe.Qdrant covers vector storage and Unsloth covers fine-tuning.ML Kit support for Android, with NPU acceleration, is coming within weeks.
Copy CodeCopiedUse a different Browserpip install -U "sentence-transformers[image,audio,video]" transformers from sentencetransformers import SentenceTransformer model = SentenceTransformer("google/embeddinggemma-2") q = model.encode("What causes the northern lights?
", promptname="SearchQuery") d = model.encode("Charged particles from the sun.", prompt_name="Document") print(model.similarity(q, d)) On Ollama, run ollama pull embeddinggemma-2.Tags range from 270m (378MB) to 740m (1.3GB).Demos live in Google AI Edge Gallery.See the developer guide for more.
Key Takeaways One 740M open model embeds text, code, images, video and audio into a shared 768d space.Modular encoders scale the footprint from 270M (text) to 740M (full multimodal).Code retrieval jumps from 68.76 to 78.68 on MTEB Code.
Runs in ~191MB to ~567MB of RAM on a Pixel 11 Pro with quantization.FAQ Can EmbeddingGemma 2 be used commercially?Yes.It is released under the Apache 2.0 license.How much memory does EmbeddingGemma 2 need?
Google reports about 191MB of active RAM for text-only use and 567MB for full multimodal use, quantized, on a Pixel 11 Pro.Check out the Model Weights on HF and Technical details.All credit goes to the researcher of this project.
Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.[Sponsored] The web is the one API most agents are missing.Databases, calendars and repos have APIs.
The open web mostly doesn’t.The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs.Search and Fetch are free.
The post Google DeepMind Releases EmbeddingGemma 2, a 740M Open Multimodal Embedding Model Built on Gemma 4 appeared first on MarkTechPost.
Related
相關文章
Reka 發表 Rho-1:一款 19B 參數的全能推理模型,能理解、生成影片並輸出機器人動作
Reka 發表了 Rho-1 的研究預覽版,這是一款從零訓練的 19B 參數全能推理模型。單一神經網路能理解與生成文字、圖片和影片,對其進行推理,並輸出機器人動作。Reka 將其定位為可直接取代透過不同模態模型傳遞工作的代理型流水線。Rho-1 的改變之處:當今多數多模態系統採用流水線架構,由中央模型規劃後,將任務分配給圖片、影片或偵測等專用模型。每一次交接都會增加延遲,且每個專用模型只能處理狹義的請求。Rho-1 消除了這些交接環節,將文字、視覺與機器人動作都轉化為單一上下文視窗中的 token。根據 Reka 的研究,一次未經編輯的會話展示了完整流程:模型繪製燈塔、標出邊框、讓它動起來、編輯影片,並最終輸出……」
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Back to Articles Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance Team Article Published October 6, 2026 Upvote 7 +1 Shaikha Alsuwaidi Shaikha710 Follow tiiuae Omar saif alkaabi Omar-Alkaabi Follow tiiuae Maitha Alhammadi MaithaAlhammadi Follow tiiuae Ahmed Alzubaidi amztheory Follow tiiuae Mohammed Alyafeai Alyafeai Follow tiiuae Leen AlQadi LeenAlQadi Follow tiiuae Basma Boussaha basma-b Follow tiiuae Hakim Hacid HakimHacid Follow tiiuae Arabic is really a family of languages living under one name.
本期AI資訊彙總2026年10月6日的產品更新、前沿研究、行業趨勢與開源項目,幫助讀者快速瞭解當天的重要動態
導航SecureDoc // EncryptionActive本期AI資訊彙總2026年10月6日的產品更新、前沿研究、行業趨勢與開源項目,幫助讀者快速瞭解當天的重要動態。
Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode
Together AI has released Together Link, a free, MIT-licensed CLI now in beta. It connects the coding agents developers already use to open models hosted on Together AI.

快手的視頻Agent,會不會來晚了?
AI價值官2026.10.05 17:36 · 來自浙江全文4219字00:00 / 11:40視頻模型廠商,正集體從"生成一段畫面"走向"交付一部成片"。文 | AI價值官,作者丨納瓦,編輯丨星野9 月 28 日,快手在模型與 Agent 兩層同時出手:白天,Agent 創作工具 likli 開啟內測;晚間,新一代模型 Kling 4.0 官宣將於 10 月上線。這並非快手的獨門動作,字節、MiniMax、阿里千問都已推出各自的創作 Agent。當模型能力被逐漸拉平,競爭正從模型轉向工作流與商業化。

AI耳機蓄勢,芯片廠商待發
半導體產業縱橫2026.10.05 15:32 · 來自內蒙古全文5079字00:00 / 15:12AI耳機還沒有成為一個邊界清晰的品類,芯片廠商卻已經開始為它準備下一代平臺。文 | 半導體產業縱橫傳統耳機,卒?AI耳機正在無線耳機市場中成為主流。據Research and Markets的測算,2025年全球AI耳機市場規模約為59.9億美元,預計到2026年將攀升至74.2億美元,並在2030年達到173.4億美元。消費電子行業裡最常被重複的一句判斷是:所有硬件,都值得用AI重新做一遍。