Evoke長程世界模型發佈
Papers arxiv:2608.
13546 Copy markdown Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Published on Aug 13 · Submitted by taesiri on Aug 14 #1 Paper of the day Upvote 82 +74 Authors: Yuanyang Yin ,Gongxuan Wang ,Yifan Zhan ,Chuanhao Li ,Kaipeng Zhang ,Feng Zhao Abstract Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency.
Generated by thinkingmachines/Inkling-Small Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model.
Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher.
Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation.
Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows.
Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons.
Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence.
A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning.
With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at 384times 640, each 1.5,s chunk is generated in 2.11,s.
As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.View arXiv page View PDF Project page Add to collection Community SII-YuanyangYin about 23 hours ago • edited about 21 hours ago Code: https://github.
com/SII-YuanyangYin/EvokePage: https://evoke-world.github.io/Evoke/YouTube Demo: https://youtu.be/QX7PBBaBGdc Reply librarian-bot about 1 hour ago This is an automated message from the Librarian Bot.I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API Wonder: Video World Model Done Better (2026) TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation (2026) AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report (2026) AlayaWorld: Long-Horizon and Playable Video World Generation (2026) Surprise Forcing: What to Remember, When to Skip in Long Video Generation (2026) OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators (2026) Generative World Renderer at the Speed of Play (2026) Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend Reply EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images Comment · Sign up or log in to comment Upvote 82 +70 Get this paper in your agent: hf papers read 2608.13546 Don't have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0 No model linking this paper Cite arxiv.org/abs/2608.
13546 in a model README.md to link it from this page.Datasets citing this paper 0 No dataset linking this paper Cite arxiv.org/abs/2608.13546 in a dataset README.md to link it from this page.Spaces citing this paper 0 No Space linking this paper Cite arxiv.org/abs/2608.13546 in a Space README.
md to link it from this page.Collections including this paper 1
Related
相關文章

消息稱字節整合 AI 生產力:TRAE、釦子併入豆包,將推統一辦公品牌“豆包工作”
作者:沁滄(實習) 責編:沁滄 評論: 感謝網友 HH_KK 的線索投遞!8 月 24 日消息,據智能湧現消息,字節跳動對旗下的辦公 AI 產品完成了一輪團隊整合:TRAE、釦子(Coze)團隊將整體併入豆包體系,其中 TRAE Work、釦子將與豆包在工作場景的產品能力進行整合;TRAE IDE 及 CLI 將作為豆包品牌下的編程產品線持續發展。

AI 大模型周榜:國產 glm-5.3-max 首秀闖入綜合榜前 15,kimi-k3-max 衝進前十
本週AI大模型Arena排行榜出現新面孔,智譜AI的glm-5.3-max首次入榜即拿下綜合榜第13名。月之暗面的kimi-k3-max排名持續攀升,成功擠進綜合榜前十,位列第10名。兩款國產大模型雙雙寫下佳績,成為本週關注焦點。

DeepSeek Harness來了:AI開始製造AI了?
DeepSeek Harness 正式推出,這項新工具被視為 AI 發展的重要里程碑,可能讓 AI 系統具備自主開發或優化其他 AI 的能力。外界關注此技術是否象徵 AI 開始「製造」AI,並可能加速人工智慧的進化與應用。目前相關細節與實際影響仍待進一步觀察。
神秘“牛來”大模型上線即登頂 背後廠商至今未揭曉
近日,一款代號為Ox Alpha的匿名AI模型在OpenRouter悄然上線,短時間內調用量迅速攀升,成功衝至平臺榜首,並刷新了該平臺的單日模型用量紀錄。不過,儘管表現驚豔,Ox Alpha背後的開發主體至今仍未揭曉。

AI辦公助手,沒有葵花寶典:五款應用萬字實測報告
AGI-Signal2026.08.24 09:12 · 來自北京全文11936字單項冠軍各有其人。2026年上半年,AI辦公賽道發生了一個根本性變化,工具不再滿足於當“對話框”,而是試圖接管完整任務,寫一段文案、做完一份報告、生成一份PPT,甚至跨應用操作。
字節整合AI辦公產品,TRAE、釦子團隊併入豆包
其中,TRAE Work、釦子將與豆包的工作場景產品能力整合;TRAE IDE及CLI則作為豆包品牌下的編程產品線繼續發展。調整後,相關產品和運營團隊統一向豆包產品負責人趙祺彙報。(iFeng Tech)TRAE與釦子此前均隸屬於字節跳動產品研發和工程架構部,前者最初定位AI編程產品,後者則聚焦AI智能體開發平臺,並持續探索不同Agent方向。