XKV讓模型傳緩存

2026年8月25日 00:00
站內 AI 整理稿

Computer Science > Artificial Intelligence arXiv:2608.

20617 (cs) [Submitted on 20 Aug 2026] Title:Dual-Cache Latent Space Communication between Heterogeneous Language Models Authors:Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang View a PDF of the paper titled Dual-Cache Latent Space Communication between Heterogeneous Language Models, by Jiyao Liu and 4 other authors View PDF HTML (experimental) Abstract:Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task.

They usually communicate by exchanging text, which puts autoregressive decoding on the critical path and reduces the exchange to a discrete message written without sight of the receiver's state.

Recent latent protocols instead translate the sharer's key-value (KV) cache into the receiver's: C2C supports heterogeneous models but requires both to read the same input, while LCF-X removes this shared-context requirement through position-free sharer-cache pooling.

Three restrictions remain: LCF-X compresses the sharer alone, supplies the same layer-local summary to every receiver position with no joint cross-layer memory to retrieve from, and assumes matched layer count and KV geometry.

We introduce XKV, which lifts all three: learned-query attention pools both caches; self-attention over receiver-aligned layer tokens, with a learned layer map reconciling different depths, mixes the pooled summaries into a compact joint memory; and a shared position decoder lets every raw receiver cache position retrieve its own per-head-gated residual in the receiver's native KV geometry.

Both models stay frozen and may differ in family, depth, KV-head count, head dimension, and tokenizer; only the translator is trained.

Across 45 dataset-model-pair settings (six heterogeneous and three same-model ordered pairings, five datasets), XKV attains the highest macro score and best average rank, improving on LCF-X on every dataset (by 4.6 exact-match and 4.

2 F1 points on ROPES) and surpassing text communication on four of the five, while training 76% fewer parameters and translating a cache pair 10.3x faster (5.8 vs.59.9 ms); end to end, XKV is 26% faster than LCF-X and 6.8x faster than text communication.Subjects: Artificial Intelligence (cs.

AI); Machine Learning (cs.LG) Cite as: arXiv:2608.20617 [cs.AI] (or arXiv:2608.20617v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.

20617 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Qi Zhang [view email] [v1] Thu, 20 Aug 2026 23:35:36 UTC (244 KB) Full-text links: Access Paper: View a PDF of the paper titled Dual-Cache Latent Space Communication between Heterogeneous Language Models, by Jiyao Liu and 4 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.

AI < prev | next > new | recent | 2026-08 Change to browse by: cs cs.LG References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

智東西生成式AI

剛剛,Codex恢復5小時限額,用戶哀嚎

AI應用風向標(公眾號:ZhidxcomAI) 作者|畢偉豪 編輯|漠影 8月25日報道,剛剛,OpenAI Codex負責人Tibo宣佈,明天開始將會恢復ChatGPT Work和Codex平臺上Plus用戶的5小時用量限制。 據Tibo所說,這一做法早就在計劃之中,目的主要是減輕算力負荷,從而在周使用額度上提供更慷慨的計劃。 這一表述暗含了計算資源不足的意思,但有網友對於Tibo的說法並不買賬,原因是前段時間他親口說過“我們不缺算力”。

剛剛
IT之家生成式AI

中國 AI 為什麼追得這麼快?美媒找到一批關鍵人物

首頁 > 智能時代>人工智能 中國 AI 為什麼追得這麼快?美媒找到一批關鍵人物 2026/8/25 12:46:15 來源:鳳凰科技 作者:於雷 責編:遠洋 評論: 8 月 25 日,據《華爾街日報》報道,中國 AI 模型近年來快速縮小與美國頭部模型的差距,其背後並非短期突然出現的技術突破,而是清華大學等高校長期形成的人才網絡,以及開源技術、人才迴流和更高算力效率等因素共同推動的結果。

剛剛
智東西生成式AI

Generalist AI融資超13億,8VC領投

機器人前瞻(公眾號:robot_pro) 作者 | 周加琦 編輯 | 漠影 機器人前瞻8月25日報道,近日,據外媒Axios報道,Generalist AI籌集了約2億美元(約13.45億元)的資金。 據消息人士透露,本次融資由8VC領投,數家現有投資者跟投。關於估值暫無具體信息,但確認高於6月份的20億美元水平。目前該公司並未回應,但在聯邦申報文件中披露了本輪融資。 根據Generalist AI向美國SEC提交的Form-D備案文件顯示,本輪計劃募資約2.08億美元(約13.99億元),實際完成融資約1.

剛剛
量子位生成式AI

人民教育音像數字出版社與小猿達成合作 “中小學課本學習智能體”首落小猿AI學習機

人民教育音像數字出版社與小猿學習機合作,推出首個「中小學課本學習智能體」,搭載於小猿AI學習機,結合權威教材與猿力大模型,提供AI伴學與AI導學雙模式。試點數據顯示,學生語文知識點掌握量提升33.6%,英語提升24%,家長陪讀時間每日減少20分鐘。雙方簽署合作協議,未來將持續推動智能體規模化應用。

剛剛
量子位生成式AI

WAIC CONNECT | AI出海,不聊虛的!CONNECT帶你拿下馬來西亞真正的AI採購需求

WAIC CONNECT MALAYSIA活動將於2026年9月7日至8日在吉隆坡舉辦,聚焦馬來西亞政府、電信、金融科技與教育等領域的實際AI採購需求,並與華為合作建立當地生態鏈接。活動規劃中國企業短講、一對一精準對接及監管機構與決策者面對面交流,僅限50至60家中國AI企業參與,旨在協助業者直接取得東南亞市場的真實商業訂單。

剛剛
鈦媒體生成式AI

被智譜封號的會員,有多少是被冤枉的?

鋅消費研究中心2026.08.25 11:33 · 來自北京全文3351字00:00 / 09:02這麼突然被封號了?文 | 鋅消費研究中心 真金白銀辦了智譜會員,結果賬號突然就被封了。 你如果遇到這種事兒,會不會賬號剛一被封,就想原地發瘋? 先別急。智譜之所以會對用戶進行封號,往往源於用戶的登錄設備、IP等出現異常,被平臺判定為多人使用賬號。 而之所以出現這種情況,除了智譜定價越來越貴之外,還源於智譜新、老會員定價不同,讓會員們有了更多套利的空間。

剛剛