XKV讓模型傳緩存
Computer Science > Artificial Intelligence arXiv:2608.
20617 (cs) [Submitted on 20 Aug 2026] Title:Dual-Cache Latent Space Communication between Heterogeneous Language Models Authors:Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang View a PDF of the paper titled Dual-Cache Latent Space Communication between Heterogeneous Language Models, by Jiyao Liu and 4 other authors View PDF HTML (experimental) Abstract:Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task.
They usually communicate by exchanging text, which puts autoregressive decoding on the critical path and reduces the exchange to a discrete message written without sight of the receiver's state.
Recent latent protocols instead translate the sharer's key-value (KV) cache into the receiver's: C2C supports heterogeneous models but requires both to read the same input, while LCF-X removes this shared-context requirement through position-free sharer-cache pooling.
Three restrictions remain: LCF-X compresses the sharer alone, supplies the same layer-local summary to every receiver position with no joint cross-layer memory to retrieve from, and assumes matched layer count and KV geometry.
We introduce XKV, which lifts all three: learned-query attention pools both caches; self-attention over receiver-aligned layer tokens, with a learned layer map reconciling different depths, mixes the pooled summaries into a compact joint memory; and a shared position decoder lets every raw receiver cache position retrieve its own per-head-gated residual in the receiver's native KV geometry.
Both models stay frozen and may differ in family, depth, KV-head count, head dimension, and tokenizer; only the translator is trained.
Across 45 dataset-model-pair settings (six heterogeneous and three same-model ordered pairings, five datasets), XKV attains the highest macro score and best average rank, improving on LCF-X on every dataset (by 4.6 exact-match and 4.
2 F1 points on ROPES) and surpassing text communication on four of the five, while training 76% fewer parameters and translating a cache pair 10.3x faster (5.8 vs.59.9 ms); end to end, XKV is 26% faster than LCF-X and 6.8x faster than text communication.Subjects: Artificial Intelligence (cs.
AI); Machine Learning (cs.LG) Cite as: arXiv:2608.20617 [cs.AI] (or arXiv:2608.20617v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.
20617 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Qi Zhang [view email] [v1] Thu, 20 Aug 2026 23:35:36 UTC (244 KB) Full-text links: Access Paper: View a PDF of the paper titled Dual-Cache Latent Space Communication between Heterogeneous Language Models, by Jiyao Liu and 4 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.
AI < prev | next > new | recent | 2026-08 Change to browse by: cs cs.LG References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

被智譜封號的會員,有多少是被冤枉的?
鋅消費研究中心2026.08.25 11:33 · 來自北京全文3351字00:00 / 09:02這麼突然被封號了?文 | 鋅消費研究中心 真金白銀辦了智譜會員,結果賬號突然就被封了。 你如果遇到這種事兒,會不會賬號剛一被封,就想原地發瘋? 先別急。智譜之所以會對用戶進行封號,往往源於用戶的登錄設備、IP等出現異常,被平臺判定為多人使用賬號。 而之所以出現這種情況,除了智譜定價越來越貴之外,還源於智譜新、老會員定價不同,讓會員們有了更多套利的空間。

中消協發佈消費提示:使用人工智能服務需謹防誤導
作者:遠洋 責編:遠洋 評論: 8 月 25 日消息,中國消費者協會今天發佈人工智能服務消費提示,提醒廣大消費者理性認識通用生成式人工智能的信息輔助和參考屬性,對涉及價格費用、合同條款、售後服務、法律責任等重要消費信息,不宜僅依據人工智能生成內容作出決定,應通過經營者官方渠道、合同文本、監管部門及其他權威渠道進行核實確認。

VC開始靠AI預測未來了
DigClaw的預測框架Rhizome在FutureX評測中,以同一套系統跨三個不同基礎模型取得前三名成績,驗證其預測能力可獨立於模型權重。該系統將搜索、因果推理與概率推斷解耦,並透過軌跡記錄與校準累積數據資產,試圖解決大型語言模型不擅長因果預測的根本問題。這項技術由AI原生創投Newborn Ventures支持,目標是將投資背後的趨勢預測系統化,應用於投資決策與產業風險評估。

全球最大生活指南百科 WikiHow 指控 OpenAI 爬取超 1.1 萬篇教程文章訓練 AI
作者:故淵 責編:故淵 評論: 感謝網友 愚公騎馬 的線索投遞!8 月 25 日消息,路透社今天(8 月 25 日)發佈博文,報道稱 WikiHow 於 8 月 21 日向美國曼哈頓聯邦法院提交訴訟,指控 OpenAI 其未經許可抓取超過 11,000 篇教程文章,用於訓練 ChatGPT 及 GPT 大型語言模型,並侵犯至少 1,200 項已註冊版權。

Edge AI Daily 早報(8月25日)
Edge AI Daily2026.08.25 08:05 · 來自北京全文4868字00:00 / 13:43微軟Skala 1.1改寫計算化學底層規則;Coherent用SiC襯底提升AI散熱25%。AI倫理調查顯示ChatGPT等模型高頻鏈接反墮胎組織。

研究團隊調查發現超七成英語生物醫學論文已使用 AI 輔助寫作,呼籲業界建立完善監管機制
作者:漾仔 責編:漾仔 評論: 8 月 25 日消息,據外媒 phys 報道,近期 Lena Holzwarth 帶頭的研究團隊分析了來自 PubMed Central 平臺的 119 萬餘篇英文生物醫學論文,發現生成式 AI 正迅速融入學術論文寫作。