VibeWorlding開源造世界

2026年9月1日 00:00
站內 AI 整理稿

2,616Curated 3D assets 323Human-annotated seed worlds 6,828Reverse-synthesized queries 59.3%Best overall Pass@1 (ours) 64.5 / 55.9Pass@1, Verified / Unverified Left: illustrative assets, size class from small to large.Right: illustrative seed 3D worlds, asset count from low to high.

How the Agent Works Each turn the agent sees the current 3D map and five rendered views, calls a tool, and the scene is re-rendered — so it can see what it built and fix it.1 Autonomous planning Infer the user's intent from the query, plan a scene layout, and decide which 3D tools to invoke.

2 Tool use via MCP Call retrieveassets, add, delete, and rotationand_translation inside the Blender sandbox.3 Observe & reflect Read back the updated 3D map plus five re-rendered views, spot mistakes, and repair them next turn.

4 Interactive 3D world Output a simulation-ready world, scored by a rubric verifier on physical feasibility and intent fulfillment.

Two task families 3D world construction builds a world from scratch given only a text query — for example, "Create a simple farm with a barn, crop fields, fences, and a few trees.

" 3D world refinement edits an existing world — for example, "Remove the car in the middle of the road and delete the two green buildings.

" The same rubric-based verifier serves double duty: it scores the leaderboard and acts as the reward service for multimodal RL post-training, so evaluation and training are measured on identical criteria.

Seed 3D Worlds Human-annotated seed worlds from VWE-Bench, spanning themes from ancient cities to desert oases.Below are six representative examples you can explore live — drag to orbit, scroll to zoom.Six representative worlds, streamed as real GLB scenes and rendered live in your browser.

Representative examples from the 323 seed worlds Rendered in the Blender sandbox, grouped by theme family.No worlds in this category.What the Agent Sees Every turn, the sandbox returns five rendered views — front, back, left, right, and top-down — alongside the textual 3D map.

This is the exact visual feedback the agent reflects on.Assets in 360° Representative examples from the 2,616-asset library, spanning the type taxonomy.Each spins automatically — drag to take control, scroll to zoom.VWE-Bench 6,828 reverse-synthesized multimodal queries across two task families.

Only asset-level edit (precise) has a ground-truth map and is rule-verified; the rest are scored by an MLLM judge against our rubrics.

Query typeSub-typeCount 3D world constructionTheme only322 Theme + elements620 Full blueprint302 Distractor120 3D world refinementAsset-level edit (precise) verified1,710 Asset-level edit (fuzzy)1,462 Scene critique553 Scene guidance757 Scene restatement477 Complex description505 Total6,828 Data splits The three splits use completely disjoint seed worlds (264 SFT / 47 RL / 12 Test, summing to 323), so leaderboard numbers measure generalization rather than seed memorization.

SplitContentsSeeds Evaluation254 cases (46 construction + 208 refinement)12 SFT5,567 cases (1,129 construction + 4,438 refinement)264 RL1,007 rollout prompts (189 construction + 818 refinement)47 Results Pass@1 on the 254-case evaluation split, reported separately on the Verified and Unverified query families.

VWE-Bench leaderboard (Pass@1).Verified (left) is rule-based and tests whether a model can hit a known target map.Unverified (right) is judged by an MLLM against intent rubrics, and is where post-training pays off most.Our VibeWorlder-30B-A3B attains the best Pass@1 on both splits (64.5 / 55.9).

Findings Frontier MLLMs are far from solving this task.Even GPT-5.5 (57.3% overall) and Qwen3.8-Max (56.9%) stay below ~60% [email protected] bottleneck traces to precise 3D world editing — hitting an exact target map, not merely producing something plausible.

Untrained open backbones fare far worse: Qwen3-VL-8B reaches 5.3% and Qwen3-VL-30B-A3B 13.6%.Post-training closes the gap to the frontier.Cold-start SFT and joint multimodal RL together lift the 30B-A3B backbone from 13.6% (base) to 34.5% (SFT) to 59.

3% after RL — the best overall Pass@1 of any model evaluated, edging out GPT-5.5.Its lead is widest on the Verified track (64.5 vs.60.4).The 8B model punches above its weight.VibeWorlder-8B reaches 41.4% overall, matching Gemini-3.1-pro (42.7%) while clearly surpassing it on Verified (59.3 vs.44.

4), exactly where rule-checkable editing is required.The two stages improve different axes.Cold-start SFT installs foundational physical and ecological competence, but delivers gains almost exclusively in 3D understanding (0.02 → 0.17).

Multimodal RL is what unlocks spatial capability: 3D understanding climbs 0.17 → 0.80 and 3D reasoning 0.04 → 0.69 — though reasoning stays below understanding, so it is only partially unlocked.Reward during RL post-training (VibeWorlder-8B).

Solid curves are the flagship run with cold-start — RL from the cold-start SFT policy; dashed curves are the w/o cold-start ablation, applying RL directly to the base backbone.

The ablation learns slowly and flattens early, while the cold-started run climbs steadily and pulls ahead on every split — most dramatically on the verified queries, where the reward more than doubles.BibTeX If you find VibeWorlding useful, please consider citing our work.

Copy citation @article{vibeworlding2026, title = {VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

}, author = {Ning, Yansong and Ye, Jingwen and Wu, Zhongkai and Sun, Yang and Zhu, Yiqin and Li, Xingyi and Zhang, Weidong and Liu, Hao}, journal = {arXiv preprint}, year = {2026} }

Related

相關文章

智東西AI Agent

Hermes新版本上線群聊模式,網友:它要幹掉馬斯克的Grok Bot了

AI應用風向標(公眾號:ZhidxcomAI) 作者|畢偉豪 編輯|漠影 “龍蝦”剛詐屍,“愛馬仕”就出手了。 9月1日報道,今天,Hermes Agent又更新了。這一次,Nous Research給它起了一個很有意思的代號:Pantheon(萬神殿)。 名字聽著很神,這次更新的內容也確實有點“眾神集結”的意思,核心在於Bot Mode的優化,這項功能雖然早早上線,但一直存在各種問題。

剛剛
智東西AI Agent

Manus宣佈恢復獨立運營

(公眾號:zhidxcom) 作者 | 程茜 編輯 | 心緣 9月1日消息,今日,爆款通用Agent產品Manus宣佈正式恢復獨立運營,Manus的三位聯合創始人CEO肖弘、首席科學家季逸超、產品合夥人張濤將繼續領導公司,並透露創新成果即將問世。

剛剛
鈦媒體AI Agent

龍蝦之父,困在了龍蝦裡

字母AI2026.09.01 12:12 · 來自北京全文6264字00:00 / 16:45OpenClaw 2.0上線,還有多少人記得它?文 | 字母AI“龍蝦”終於有了新動靜,但“龍蝦熱”早已過去。8月31日,OpenClaw發佈了v2026.8.1版本,官方將它稱之為“OpenClaw 2.0”。再也沒有以前那麼高的使用門檻了,而且OpenClaw 2.0還加入了大量新功能,比如整理記憶、多人協作等等。

53 分鐘前

這款Agent,想做千萬畢業生的“求職搭子” | 水下項目

求職市場的焦慮從未像這幾年一樣具體而迫切。每年上千萬畢業生湧入人力市場,有人海投數百封履歷卻杳無音訊,有人好不容易進入面試卻在最後一關被刷下;而企業端同樣頭痛,HR被大量格式千奇百怪的履歷淹沒,初篩一個基層職位可能要同時比對上百位候選人。就在供需雙方都被資訊鴻溝與重複勞動困住之際,一款主打「求職搭子」定位的Agent產品悄然浮出水面,試圖用AI代理的方式,替年輕人承接這條漫長求職路上一件件瑣碎、磨人卻又至關重要的小事。

2 小時前
雷峰網AI Agent

從生成工具到專業創作夥伴,小云雀AI官宣品牌升級

8月31日,內容創作平臺小云雀AI宣佈品牌升級,圍繞“一起創作好故事”這一全新定位,升級產品能力、創作者扶持政策和內容生態。未來3年,小云雀AI將投入一億積分,依託創作者計劃、劇本大賽等活動,支持更多創作者講出好故事。小云雀Slogan更新全鏈路產品能力升級,支持完整故事創作4月上線以來,小云雀創作者計劃吸引眾多創作者參與,並湧現出《喪屍清道夫》《餘燼之後》《歸墟》等多部爆款作品。截至目前,創作者通過小云雀製作的AI短劇全網累計播放量超過60億次,並誕生超過30部千萬播放級AI故事短片。

3 小時前