訂房代理的價格偏好被量化對比
Economics > General Economics arXiv:2609.
31468 (econ) [Submitted on 25 Sep 2026] Title:PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents Authors:Pavel Kireyev View a PDF of the paper titled PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents, by Pavel Kireyev View PDF HTML (experimental) Abstract:LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs.
Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences.
We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties.
We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently.
What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks.
What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.Comments: Accepted to EMNLP 2026 Industry Track.19 pages, 10 figures, 6 tables.Code and data: this https URL Subjects: General Economics (econ.
GN); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2609.31468 [econ.GN] (or arXiv:2609.31468v1 [econ.GN] for this version) https://doi.org/10.48550/arXiv.2609.
31468 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Pavel Kireyev [view email] [v1] Fri, 25 Sep 2026 16:17:28 UTC (2,036 KB) Full-text links: Access Paper: View a PDF of the paper titled PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents, by Pavel KireyevView PDFHTML (experimental)TeX Source view license Current browse context: econ.
GN < prev | next > new | recent | 2026-09 Change to browse by: cs cs.AI cs.CL econ q-fin q-fin.EC References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

Claude Sonnet 5.5發佈:性能逼近Opus 5.5,價格減半,安全防護升級
更快、更便宜了。在推出旗艦級模型Claude Opus 5.5僅一週後,Anthropic又帶來了Claude 5.5家族的第二款產品。當地時間9月28日,Anthropic正式發佈Claude Sonnet 5.5。作為面向日常辦公、軟件開發等場景的主力模型,Sonnet 5.

Manus這次,直接衝著“人”去了
字母AI2026.09.29 16:47 · 來自北京全文4135字00:00 / 10:57一個Personal Agent 不過癮,Manus要配齊“秘書班子”。文 | 字母AIMuse剛把個人Agent推到聚光燈下,Manus就帶著它的“Agent團隊”回來了。

OpenAI因新模型太強叫停發佈
AGI計劃暫停。 程淺 發自 凹非寺 | 公眾號 QbitAI 原定於10月亮相的GPT-6.1 Astra突然被曝暫停發佈! 也就是說,OpenAI,三天連踩兩腳剎車。 而因為幾天前的DNS事件,OpenAI同時暫停了內部最強模型所有涉及工具調用的訓練、評測和推理。 關於此事件可以參考:啥題啊能幹崩OpenAI最強模型訓練。該事件被官方稱作Hugging face之後的首次模型失調事件。 這次暫停指向了模型訓練中的一個越發棘手的問題。

OpenAI明日重啟200美元Pro訂閱:API配額減半,取消5小時限制
這一策略調整標誌著其在大模型商業化路徑上的計費邏輯發生實質性轉變。官方說明指出,新版訂閱徹底解除了過往備受爭議的5小時使用時長限制,讓開發者與企業用戶得以完全自由支配每週配額。配額表面減半的背後,實質是底層模型推理效率的大幅躍升與API定價的階梯式下調。

AMD 82 億美元全股票吞下World Labs,李飛飛掛帥執行副總裁
更吸睛的是人——World Labs 聯合創始人、有著 AI 教母之稱的李飛飛,將加入 AMD 出任執行副總裁兼首席科學家,直接向蘇姿豐彙報。World Labs 是李飛飛等人於2024年聯合創辦的,主攻的方向叫世界模型:讓 AI 根據文本、圖像和視頻,生成或重建出能夠交互的3D 環境。

成立一年完成5輪融資,諾因智能再獲數億元,累計超10億元
諾因智能宣布完成數億元人民幣天使+++輪融資,由京東相關基金領投,累計融資額超過10億元人民幣。資金將用於擴充訓練數據、推進KNOWIN-X1本體量產準備並延攬人才。公司預計2027年第一季從技術進展走向消費市場,將具身大模型GLOW的「一教就會」能力帶入真實家庭。