Evoke長程世界模型發佈
Papers arxiv:2608.
13546 Copy markdown Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Published on Aug 13 · Submitted by taesiri on Aug 14 #1 Paper of the day Upvote 82 +74 Authors: Yuanyang Yin ,Gongxuan Wang ,Yifan Zhan ,Chuanhao Li ,Kaipeng Zhang ,Feng Zhao Abstract Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency.
Generated by thinkingmachines/Inkling-Small Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model.
Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher.
Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation.
Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows.
Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons.
Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence.
A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning.
With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at 384times 640, each 1.5,s chunk is generated in 2.11,s.
As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.View arXiv page View PDF Project page Add to collection Community SII-YuanyangYin about 23 hours ago • edited about 21 hours ago Code: https://github.
com/SII-YuanyangYin/EvokePage: https://evoke-world.github.io/Evoke/YouTube Demo: https://youtu.be/QX7PBBaBGdc Reply librarian-bot about 1 hour ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API Wonder: Video World Model Done Better (2026) TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation (2026) AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report (2026) AlayaWorld: Long-Horizon and Playable Video World Generation (2026) Surprise Forcing: What to Remember, When to Skip in Long Video Generation (2026) OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators (2026) Generative World Renderer at the Speed of Play (2026) Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 82 +70 Get this paper in your agent: hf papers read 2608.13546 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0 No model linking this paper Cite arxiv.org/abs/2608.
13546 in a model README.md to link it from this page.Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.13546 in a dataset README.md to link it from this page.Spaces citing this paper 0 No Space linking this paper Cite arxiv.org/abs/2608.13546 in a Space README.
md to link it from this page.Collections including this paper 1
Related
相關文章

AI辦公助手,沒有葵花寶典:五款應用萬字實測報告
AGI-Signal2026.08.24 09:12 · 來自北京全文11936字單項冠軍各有其人。2026年上半年,AI辦公賽道發生了一個根本性變化,工具不再滿足於當“對話框”,而是試圖接管完整任務,寫一段文案、做完一份報告、生成一份PPT,甚至跨應用操作。

Anthropic新模型偷「吃瓜」,最強Fable 5爆冷
Anthropic 近日推出新款 AI 模型,在內部測試中意外展現「吃瓜」能力,引發社群熱議。該模型不僅能快速理解網路迷因與流行語,更在特定任務上表現出人意料,讓原本被外界視為最強對手的 Fable 5 爆冷落後,業界對這項結果感到相當驚訝。目前 Anthropic 官方尚未針對模型實際表現與測試細節做出完整說明,市場則持續關注後續可能的技術更新與應用方向。
端側AI大洗牌:vivo藍心登頂手機大模型榜首,3B小參數跑分逼近雲端巨頭
SuperCLUE發佈手機端側大模型測評,vivo藍心BlueLM3.5 Nano 3B以89.86分居綜合第一。該3B小模型得分逼近谷歌Gemini3.6 Flash、豆包Seed2.1 Pro、千問Qwen3.8 Max等雲端大模型,展現端側性能突破。
Kimi K2.5 月底退役:月之暗面第一代萬億參數多模態模型謝幕
月之暗面官宣第一代萬億參數多模態模型Kimi K2.5將於本月底結束服役。該模型今年1月推出並開源,是Kimi迄今最全能模型,採用原生多模態架構,支持視覺與文本輸入、思考/非思考模式、對話與Agent任務,在Agent、代碼、圖像、視頻及通用智能取得開源SOTA。K3將接力,參數規模再上臺階。
光子躍遷亮相BIRTV 2026:以"AI+影像"重構創作範式,三大板塊解碼下一代影像生態
8月19日,BIRTV 2026(北京國際廣播電影電視展覽會)在北京拉開帷幕。在這場匯聚全球廣電與影像領域頂尖技術與創意的盛會上,光子躍遷以"AI+影像"為核心敘事,攜個人智能影像生態重磅亮相,向行業展示了一個由AI驅動、以人為中心的影像未來。與行業展會常見的深色科技風不同,光子躍遷的展臺以純淨白色為主基調,輔以品牌藍色進行點睛點綴,在千篇一律的深色展臺中脫穎而出,傳遞出品牌年輕、活力、面向未來的基因。

諾亦騰機器人發佈 HiPHI,開源 617.5 小時高精度人體運動數據
作者:潞源 責編:潞源 評論: 8 月 23 日消息,諾亦騰機器人在 2026 世界機器人大會期間發佈 HiPHI,這是一套面向人形機器人學習、數字人,以及計算機圖形學領域研究人員和工程師的高精度光學動作捕捉數據集。據報道,該數據集總長 617.