何夕2077生成式AI

浙大提出評估框架

2026年8月9日 00:00

重點摘要

浙江大學團隊提出ProVisE評估框架與SpatialGen-Bench基準,讓影像生成模型在像素空間直接輸出空間判斷,解決現有基準與模型輸出格式不匹配的問題。實驗顯示,影像生成模型在可直接用像素表達空間答案的任務上表現出色,而文字輸出模型在組合式空間推理上仍具優勢。

站內 AI 整理稿

Computer Science > Computer Vision and Pattern Recognition arXiv:2607.

21072 (cs) [Submitted on 23 Jul 2026] Title:Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Authors:Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang View a PDF of the paper titled Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text, by Xu Wang and 6 other authors View PDF HTML (experimental) Abstract:Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world.

Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols.

Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models.

This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space.

We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics.

ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks.We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms.

We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks.

Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning.

These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.Comments: 36 pages, 14 figures.

Project page: this https URL Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.21072 [cs.CV] (or arXiv:2607.21072v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.

21072 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Xu Wang [view email] [v1] Thu, 23 Jul 2026 09:04:48 UTC (26,111 KB) Full-text links: Access Paper: View a PDF of the paper titled Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text, by Xu Wang and 6 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.

CV < prev | next > new | recent | 2026-07 Change to browse by: cs References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

IT之家生成式AI

首個全國產 10 萬卡 AI 超集群投用:每秒峰值算力相當全人類持續計算 200 年

首頁 > 智能時代>人工智能 首個全國產 10 萬卡 AI 超集群投用:每秒峰值算力相當全人類持續計算 200 年 2026/8/9 12:23:34 來源:IT之家 作者:浩渺 責編:浩渺 評論: IT之家 8 月 9 日消息,據央視新聞今日報道,記者從國家發展改革委瞭解到,今年以來,我國算力底座進一步夯實,首個全國產 10 萬卡人工智能超集群日前正式投用。

剛剛
量子位生成式AI

當題庫追不上模型,AI開始給自己出題:中國這支團隊跑通了數據層RSI

中國無盡前沿團隊發布BigBang-V1基座模型,這是首個透過遞歸自我改進(RSI)原生方式訓練的模型,參數規模35B,在某些科研任務上表現超越1T級別模型。該模型的訓練數據百分之百由AI自主合成,透過可驗證前沿任務讓系統持續生成高品質數據,實現了數據層的自我進化閉環,不需人類逐題參與。

剛剛
量子位生成式AI

Opus 5狂燒6.9億token做遊戲,GPT-5.6用5美元復刻了

Opus 5 使用6.9億個token和一條長達2000字的提示詞,花費423美元生成了一款美式卡通風格的水上快艇競速遊戲《INK TIDE》,品質精美。隨後有網友用GPT-5.6在Codex中以約5美元成本快速復刻出類似遊戲,但完成度與畫面細節明顯較差。此案例顯示當前AI生成遊戲的品質上限仍取決於提示詞設計與預算投入。

剛剛
量子位生成式AI

爆料:哈薩比斯原本要和Jeff Dean一起走!

據報導,哈薩比斯原本計畫與Jeff Dean等人同期離開谷歌,但谷歌管理層擔心股價受挫而極力挽留,最終他卸任Google DeepMind CEO,改任董事長及Alphabet首席科學家,被外界視為緩兵之計,可能一年內徹底離開。與此同時,谷歌已將DeepMind的管理權、關鍵人才與決策中心逐步移往加州,而哈薩比斯未來可能更深入投入Isomorphic Labs,延續AI加速藥物發現的科學願景。

剛剛
鈦媒體生成式AI

Jeff Dean離職谷歌首次亮相回答:AI的下一個十年是什麼

前Google首席科學家Jeff Dean離職後首次公開亮相,在AASF 2026峰會上闡述他對AI下一個十年的構想,並發布新公司Discovery Loop。他強調技術判斷應以「10倍提升」為標準,並認為雲端運算將基礎設施變成商品,讓頂尖人才得以在小型創業團隊中追求前沿突破。

剛剛
IT之家生成式AI

消息稱小米成立具身智能與應用部,前字節 Seed 具身負責人孔濤掛帥

小米機器人事業部進行大調整,新成立具身智能與應用部,由前字節跳動 Seed Robotics 具身智能負責人孔濤主導。該部門整合了多個機器人團隊,孔濤將負責小米所有具身智慧業務。此外,小米機器人在汽車工廠的實習表現優異,自攻螺母上件工站雙側作業成功率已提升至98%,接近人類水準。

剛剛