Cohere 發布 Embed 5:與 Voyage 4 Large、Gemini Embedding 2 及 OpenAI 的比較
Cohere has released Embed 5, a new embedding model family.It targets enterprise search, RAG, and agentic retrieval.The model family ships in 2 tiers.Embed 5 Pro targets maximum retrieval quality.Embed 5 Fast targets latency and cost on the live query path.
Both accept text, images, and fused text plus image inputs.Both cover 100+ languages and read up to 128K tokens.The key design choice: Pro and Fast share 1 embedding space.You can index with one and query with the other.Is it deployable today?
Yes, both tiers are generally available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker.Private VPC or on-prem serving runs through vLLM.What Cohere Shipped The API model IDs are embed-v5.0-pro and embed-v5.0-fast, per Cohere’s model docs.
Both output 2048, 1536, 1024, 768, 512, or 256 dimensions, with 2048 as default.Embeddings come back as float, int8, or binary.Pro costs $0.12 per 1M text tokens.Fast costs $0.08.Image inputs cost $0.40 per 1M tokens on both.Embed 5 can embed a page image directly.
It can also fuse an image with its metadata into a single vector.That is important for scanned pages, slide decks, schematics, and charts, where text extraction drops information.Pro and Fast: One Index, Two Query Paths Cohere tested every corpus and query pairing across 40 development datasets.
Normalized to Pro plus Pro at 100, a Pro index queried with Fast scored 98.4.An all-Fast setup scored 96.6.Cohere’s recommended pattern is to index with Pro and query with Fast.One constraint: both sides must use the same output dimension.The split targets agentic workloads.
An agent may issue dozens of searches per task, and query latency compounds.Cohere team reports Fast processed 377.3 documents per second versus 159.7 for Pro.Benchmarks On ViDoRe V3, Embed 5 Pro averages 85.8, an 8.8-point gain over Embed 4.Fast averages 84.5.Voyage 4 Large scores 83.
7, Gemini Embedding 2 scores 83.2, and OpenAI text-embedding-3-large scores 75.5.On Cohere’s parsed-PDF suite, Pro leads at 84.8 against Voyage 4 Large at 83.6.Finance is the strongest showing.Pro ranks first on FinanceBench (80.1), FinQA (90.0), and ViDoRe V3 Finance (85.0).
Fast ranks second on all 3.Multilingual results are mixed.Pro leads the European-language average at 77.However, Gemini Embedding 2 beats Pro on 9 of 10 further languages in Cohere’s own results table.Those include Japanese, Arabic, Hindi, and Telugu.One important thing to note.
Most numbers use RCP-nDCG@10, a new Cohere metric.It reorders a fixed candidate set, so it measures reranking quality more than first-stage retrieval.Cohere published the evaluation code, but independent replication is still pending.
Storage Costs at Scale Embed 5 uses Matryoshka representation learning plus lower-precision outputs.A 2048-dim float32 vector needs 8 KB.A 1024-dim int8 vector needs 1 KB.A 256-dim binary vector needs 32 bytes.Across 100M chunks, raw storage drops from about 819 GB to 3.2 GB.
Cohere recommends 1024-dim int8 as the default, citing near-full-precision quality.
Interactive Explainer: How Embed 5 Works Cohere Embed 5 Explainer | Marktechpost #mtp-ce5{--bg:#06080F;--panel:#0D1322;--panel2:#121A2E;--line:#1E2A45;--blue:#3E7BFA;--cyan:#5CC8FF;--coral:#FF7759;--txt:#E8EEFC;--mut:#8C9BBA;--mtp:#76B900;font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Inter,Roboto,Arial,sans-serif;background:var(--bg)!
important;color:var(--txt)!important;border:1px solid var(--line)!important;border-radius:16px;padding:22px;max-width:860px;margin:0 auto;box-sizing:border-box;overflow:hidden;position:relative} #mtp-ce5 *{box-sizing:border-box} #mtp-ce5 p:empty,#mtp-ce5 del,#mtp-ce5 s{display:none!
important} #mtp-ce5 .hd{display:flex;justify-content:space-between;align-items:center;gap:12px;margin-bottom:14px} #mtp-ce5 .tag{font-size:11px;letter-spacing:.14em;text-transform:uppercase;color:var(--cyan)!important;font-weight:700} #mtp-ce5 .cnt{font-size:12px;color:var(--mut)!
important;font-variant-numeric:tabular-nums} #mtp-ce5 .bar-top{height:3px!important;background:var(--line)!important;border-radius:3px;overflow:hidden;margin-bottom:18px} #mtp-ce5 .bar-top i{display:block;height:3px;background:linear-gradient(90deg,var(--blue),var(--cyan))!
important;transition:width .5s ease} #mtp-ce5 .slide{display:none} #mtp-ce5 .slide.on{display:block;animation:ceIn .45s ease} @keyframes ceIn{from{opacity:0;transform:translateY(10px)}to{opacity:1;transform:none}} #mtp-ce5 h3{font-size:22px;line-height:1.25;margin:0 0 6px;color:var(--txt)!
important;font-weight:800} #mtp-ce5 .sub{font-size:14px;line-height:1.55;color:var(--mut)!important;margin:0 0 16px} #mtp-ce5 .card{background:var(--panel)!important;border:1px solid var(--line)!important;border-radius:12px;padding:14px} #mtp-ce5 .row{display:flex;gap:12px;flex-wrap:wrap} #mtp-ce5 .
btn{background:var(--panel2)!important;color:var(--txt)!important;border:1px solid var(--line)!important;border-radius:999px;padding:8px 14px;font-size:13px;font-weight:600;cursor:pointer;transition:all .2s;font-family:inherit} #mtp-ce5 .btn:hover{border-color:var(--blue)!important} #mtp-ce5 .btn.
act{background:var(--blue)!important;border-color:var(--blue)!important;color:#fff!important;box-shadow:0 0 18px rgba(62,123,250,.45)} #mtp-ce5 .go{background:linear-gradient(90deg,var(--blue),var(--cyan))!important;color:#04101F!important;border:0!important} #mtp-ce5 .
lbl{font-size:11px;letter-spacing:.1em;text-transform:uppercase;color:var(--mut)!important;margin:0 0 8px;font-weight:700} #mtp-ce5 .q{font-size:14px;color:var(--txt)!important;background:var(--panel2)!important;border:1px dashed var(--blue)!
important;border-radius:10px;padding:10px 12px;margin-bottom:12px} #mtp-ce5 .vec{display:flex;gap:3px;height:54px;align-items:flex-end;margin:10px 0 14px} #mtp-ce5 .vec span{flex:1;background:linear-gradient(180deg,var(--cyan),var(--blue))!
important;border-radius:2px 2px 0 0;height:4%;transition:height .6s cubic-bezier(.2,.8,.2,1)} #mtp-ce5 .doc{display:flex;align-items:center;gap:10px;padding:9px 10px;border:1px solid var(--line)!important;border-radius:10px;margin-bottom:8px;background:var(--panel2)!important;transition:all .
4s} #mtp-ce5 .doc.win{border-color:var(--cyan)!important;box-shadow:0 0 0 1px var(--cyan) inset,0 0 20px rgba(92,200,255,.25)} #mtp-ce5 .doc .t{flex:1;font-size:13px;line-height:1.4;color:var(--txt)!important} #mtp-ce5 .doc .m{width:90px;height:8px!important;background:var(--line)!
important;border-radius:8px;overflow:hidden} #mtp-ce5 .doc .m i{display:block;height:8px;width:0;background:var(--cyan)!important;transition:width .9s ease} #mtp-ce5 .doc .v{width:38px;text-align:right;font-size:12px;font-variant-numeric:tabular-nums;color:var(--cyan)!important} #mtp-ce5 .
note{font-size:11.5px;color:var(--mut)!important;margin-top:10px;line-height:1.5} #mtp-ce5 .grid2{display:grid;grid-template-columns:1fr 1fr;gap:12px} #mtp-ce5 .gauge{position:relative;width:170px;height:170px;margin:4px auto 8px} #mtp-ce5 .gauge svg{transform:rotate(-90deg)} #mtp-ce5 .gauge .
val{position:absolute;inset:0;display:flex;flex-direction:column;align-items:center;justify-content:center} #mtp-ce5 .gauge .val b{font-size:34px;font-weight:800;color:var(--txt)!important;font-variant-numeric:tabular-nums} #mtp-ce5 .gauge .val small{font-size:11px;color:var(--mut)!
important} #mtp-ce5 .hb{margin:10px 0} #mtp-ce5 .hb .n{display:flex;justify-content:space-between;font-size:12.5px;margin-bottom:5px;color:var(--txt)!important} #mtp-ce5 .hb .n em{font-style:normal;color:var(--mut)!important;font-variant-numeric:tabular-nums} #mtp-ce5 .hb .tr{height:12px!
important;background:var(--panel2)!important;border-radius:6px;overflow:hidden;border:1px solid var(--line)!important} #mtp-ce5 .hb .tr i{display:block;height:100%;width:0;border-radius:6px;background:linear-gradient(90deg,var(--blue),var(--cyan))!important;transition:width 1s cubic-bezier(.2,.8,.
2,1)} #mtp-ce5 .hb.other .tr i{background:#3A4766!important} #m
Related
相關文章

何愷明團隊新作:看貓片就能學會ARC挑戰
何愷明團隊提出NAT-ARC,一種純視覺的ARC解題方案,不依賴語言模型,而是使用ImageNet上的自然圖像進行MAE預訓練,再遷移到抽象格子推理任務。該方法在ARC-1上達到63.4%的單模型pass@2分數,集成後提升至70.2%,逼近專用LLM系統的表現,並證明了視覺預訓練能有效破解抽象推理的scaling瓶頸。

谷歌Gemini 4突然發佈!RSI加持,GPT和Opus都讓讓
。。。 假期第一天,谷歌攜Gemini 4 Argon空降多榜單第一。 拳打Opus 5.5,腳踢GPT-6 Astra。 更誇張的還在後頭,單項任務成本最低可至1.99美元,直接是Astra費用砍半。 最高百萬Token輸出上限,面向編程、金融和法律等複雜工作流,而且劃重點,網絡安全防禦能力超牛掰。 這波等等黨要贏麻了。 不過吧,咱普通用戶現在只可遠觀,暫時還吃不上。

階躍星辰“第一梯隊”,是“自嗨”嗎?
AIX財經2026.10.01 18:22 · 來自福建全文5252字00:00 / 15:27同行各自跑出了主線,階躍的“全棧”能跑通嗎?文 | AIX財經,作者|雷晶,編輯|魏佳沉寂許久的階躍星辰,正試圖擠進大模型第一梯隊。9月20日,它發佈Step 5 Preview,原生支持文本與視覺輸入,重點面向編程、軟件工程、專業知識工作和長程Agent任務。上線初期,登錄並完成首次調用後可獲得30天的Step Plan免費使用權。幾天後,Step Plan的月度套餐一度售罄。

騰訊經銷、字節駐場、Kimi借船:FDE成了大模型的新成本?
新立場Pro2026.10.01 16:16 · 來自四川全文5421字00:00 / 16:08AI公司活成了它們最討厭的樣子。文 | 新立場Pro今年 3 月 6 日上午十點,深圳騰訊大廈樓下開始排隊。人們帶著電腦,等騰訊雲工程師幫自己安裝 OpenClaw,首批八十多人在十點開始排隊,到十一點,數百個預約號已經發完。

谷歌推出了個“做題家”:Gemini 4 Argon屠榜,但幹活差點意思
字母AI2026.10.01 16:16 · 來自北京全文3510字00:00 / 09:31消失10個月的谷歌,為什麼連Pro這塊招牌都扔了?文 | 字母AIGemini 4可算來了,連名字也換了:這次的旗艦不叫Pro,叫Argon。從谷歌公佈的成績看,Gemini 4 Argon在知識工作方面表現搶眼,多項測試超過了OpenAI和Anthropic的旗艦模型。它主打一個知識面廣,從金融、法律到數學、科學,都有不錯的表現。谷歌還確認,9月15日在LMArena上亮相、表現接近GPT-6 Astra的“3.

豪擲82億美元,芯片巨頭押注的世界模型究竟是什麼
FoST未來敘事2026.10.01 16:04 · 來自北京全文3694字00:00 / 10:23當AI開始生成“世界”,影音遊內容正在被顛覆。文 | FoST未來敘事,作者 | 蔥蔥82億美元,芯片巨頭AMD以全股票形式收購世界模型公司World Labs。AMD在收購公告裡寫道:“World Labs 的模型經驗,將幫助AMD更深入地理解工作負載如何演變,進而塑造未來的技術路線圖。”換句話說,AMD看中的,不只是World Labs今天的模型能力,更是世界模型這類前沿模型可能給芯片硬件研發帶來的“前瞻性”。