浙大提出評估框架
Computer Science > Computer Vision and Pattern Recognition arXiv:2607.
21072 (cs) [Submitted on 23 Jul 2026] Title:Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Authors:Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang View a PDF of the paper titled Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text, by Xu Wang and 6 other authors View PDF HTML (experimental) Abstract:Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world.
Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols.
Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models.
This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space.
We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics.
ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks.We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms.
We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks.
Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning.
These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.Comments: 36 pages, 14 figures.
Project page: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.21072 [cs.CV] (or arXiv:2607.21072v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.
21072 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Xu Wang [view email] [v1] Thu, 23 Jul 2026 09:04:48 UTC (26,111 KB) Full-text links: Access Paper: View a PDF of the paper titled Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text, by Xu Wang and 6 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.
CV < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

華為大模型雙子星聯手造物理基模,剛融資數億元
機器人前瞻(公眾號:robot_pro) 作者 | 許麗思 編輯 | 漠影 9月23日報道,近日,物理AI公司息壤開物(XIRRA)已在兩個月內連續完成數億元種子輪及天使輪融資,由敦鴻資產領投,華控基金、三花控股、銀杏谷資本、澤然資本、本堅基金、上海天使會、標樸投資等知名機構跟投。

Jev 的「百億補貼」迷局:既然「極省」為何還要狂送 1.2 億 Token?
本文作者: 鄭佳美 2026-09-23 12:10 導語:先壓低判斷成本,再用 Calibration 提升可靠性,最終借 Agent 放大調用量。先壓低判斷成本,再用 Calibration 提升可靠性,最終借 Agent 放大調用量。

豆包做手機、阿里造平板:AI開始爭奪硬件控制權
識礁Farsight2026.09.23 11:43 · 來自北京全文3182字00:00 / 09:20大廠為何突然集體“變硬”?文 | 識礁Farsight繼豆包推出“第二代豆包手機”後,阿里也開始“變硬”。2026年9月22日,阿里召開2026雲棲大會,展示了仍在研發中的“千問平板”QwenBook。據《晚點LatePost》報道,QwenBook由阿里雲旗下無影團隊打造,定位原生智能體電腦,產品定義和硬件部分由團隊自研。無影過去做雲電腦積累的系統和端雲能力,將融入其中。

螞蟻密算韋韜:企業AI進不了生產,卡點不在模型
TechPulse2026.09.23 11:40 · 來自浙江全文1839字00:00 / 05:57韋韜提到,用自然語言描述不許做什麼這類紅線時,大語言模型的違背率超過80%。企業AI的下一道門檻,落在了一個比模型能力更難工程化的問題上。

韶音首款 AI 耳機 OpenFit 2 AI 亮相雲棲大會,重磅首發“AI 實驗室”四大新功能
作為開放式耳機領域的領頭羊,Shokz 韶音首次重磅參展,不僅攜旗下首款 AI 耳機 OpenFit2AI 核心亮相,更在現場聯合光帆科技首發了“AI 實驗室”全新功能,全面展現了 Qwen 大模型在音頻可穿戴設備中的深度產業落地。作為韶音佈局智能音頻賽道的開山之作,OpenFit2AI 深度融合了通義千問(Qwen)大模型的強大能力,集長時錄音、AI 會議紀要生成與多語種實時翻譯於一體。

GPT-6 Sol和Luna上線,打折比梁文鋒還狠
字母AI2026.09.23 11:08 · 來自北京全文5453字00:00 / 14:27Opus 5.5發佈90分鐘,OpenAI來打價格戰了。文 | 字母AI今夜註定無眠。Opus 5.5剛剛發佈90分鐘,OpenAI就把GPT-6 Sol和Luna一起端了上來。前腳Anthropic剛把Opus的價格砍了一刀,後腳OpenAI直接把刀砍得更深。GPT-6 Sol每百萬Token輸入只要2美元、輸出10美元,價格正好只有Opus 5.5的一半,Astra的五分之一。Luna更誇張,只要0.1美元和0.