訂房代理的價格偏好被量化對比
Economics > General Economics arXiv:2609.
31468 (econ) [Submitted on 25 Sep 2026] Title:PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents Authors:Pavel Kireyev View a PDF of the paper titled PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents, by Pavel Kireyev View PDF HTML (experimental) Abstract:LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs.
Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences.
We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties.
We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently.
What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks.
What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.Comments: Accepted to EMNLP 2026 Industry Track.19 pages, 10 figures, 6 tables.Code and data: this https URL Subjects: General Economics (econ.
GN); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2609.31468 [econ.GN] (or arXiv:2609.31468v1 [econ.GN] for this version) https://doi.org/10.48550/arXiv.2609.
31468 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Pavel Kireyev [view email] [v1] Fri, 25 Sep 2026 16:17:28 UTC (2,036 KB) Full-text links: Access Paper: View a PDF of the paper titled PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents, by Pavel KireyevView PDFHTML (experimental)TeX Source view license Current browse context: econ.
GN < prev | next > new | recent | 2026-09 Change to browse by: cs cs.AI cs.CL econ q-fin q-fin.EC References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...
Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.
ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?
) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.
AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?
) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)
Related
相關文章

82億美元,硅谷上演頂級版“Girls help girls”
82億美元,硅谷上演頂級版“Girls help girls”山農下山2026.09.29 18:12 · 來自甘肅全文4164字蘇姿豐與李飛飛“會師”。 文 | 山農下山溫柔的收購與現實的考量“去尋找一個更新的世界,永遠不會太晚。”AMD在9月28日官宣收購李飛飛創辦的 World Labs,82億美元、全股票交易。 隨後,李飛飛在個人長信中引用了英國詩人丁尼生在《尤利西斯》的這句話。48歲創業,2年後將公司以82億美金的價格賣給AMD,成為後者歷史上的第二大收購案。交易完成後,李飛飛出任AMD執行副總裁兼首席科學家,直接向董事長兼CEO蘇姿豐彙報。這句話是中年人給自己的鼓勵,更是寫給蘇姿豐與AI行業的投名狀。蘇姿豐是World Labs的早期投資人。李飛飛在官宣長文裡稱她為“great friend”。今年1月的CES,蘇姿豐把李飛飛請上主題演講的舞臺,現場演示她的產品Marble在AMD芯片上跑出四倍加速;更早之前,兩家公司已經開始聯合訓練模型。回頭來看,兩人的關係是步步升溫,走到收購這一步,似乎也是水到渠成。硅谷見過無數劍拔弩張的收購戰:甲骨文與仁科、微軟與雅虎、博通與高通,每一段都刺激到值得用一本書去記錄。與之相比,AMD 對 Wordl Labs 的這場收購,堪稱溫柔。考慮到蘇姿豐與李飛飛同為華裔女性的身份背景,這也堪稱硅谷頂級的“Girls help girls”。站在收購雙方的兩位女性,都是硅谷舉足輕重的人物了。蘇姿豐在2014年接手AMD,帶領這家當時市值20多億美元、賬上全是虧損的老牌公司實現復興,12年後市值一度突破萬億美元。其中關鍵就在於,工程師出身的蘇姿豐先後押準了Zen架構、AI技術。在面臨複雜且宏大的局面時,她很擅長先縮小戰線,再圍繞一個確定性的技術支點,去建立重構。眼下,物理AI正在成為AI產業下一階段的重要探索方向。圍繞大語言模型的競爭,似

安卓端Google Assistant正式謝幕,Gemini接管手機語音入口,手錶汽車卻還留著舊管家
AI資訊AI新聞資訊正文安卓端Google Assistant正式謝幕,Gemini接管手機語音入口,手錶汽車卻還留著舊管家發佈於AI新聞資訊發佈時間 :2026年9月29號 17:59閱讀 :1分鐘谷歌對自家老牌語音助手的退休程序,終於走到了無法回頭的那一步。9月29日據外媒報道,谷歌從上月起就開始向用戶群發郵件,預告手機端的 Google Assistant 即將停用;如今大批安卓用戶反饋,他們的手機已經實實在在調不動那個熟悉的語音助手了。一些用戶更發現,這幾天 Gemini 應用的賬戶菜單裡,切換到 Google Assistant 的入口已經消失——從選項到服務,舊助手正在被一點點抹掉。這場替代並不讓人意外。谷歌把 Gemini 推上主位已是公開的策略,但落到真實用戶身上,感受卻五味雜陳。不少人覺得 Google Assistant 本來就把需求伺候得挺好,而 Gemini 至今仍帶著幻覺等老毛病,還少了部分離線語音指令——也就是說,新管家更聰明,卻在某些時刻不如舊管家隨手、可靠。值得劃清邊界的是,這次變動只掃到移動設備和手機配件這一層:智能手錶、Android Auto 車機、耳機裡的 Google Assistant 一併被納入停用範圍;而智能音箱和智能顯示器這邊,谷歌此前明確表示仍會繼續支持 Google Assistant。換句話說,家裡客廳那隻會播報天氣的音箱暫時不用換口音,手機口袋裡的那個,卻已經被 Gemini 正式接管。當一代陪伴了無數安卓用戶起步的語音助手在手機上熄燈,谷歌押注的顯然是同一個入口裡長出更強大腦的未來——只是對習慣了離線喊一嗓子就能定鬧鐘的人而言,這場交接還帶著一點沒磨平的毛刺。相關推薦谷歌在印度測試 Gemini 內直購 Flipkart 商品,全面進軍端內交易生態谷歌正加速推動AI服務從“商品發現”邁向“交易落地”。最新消息稱,谷

Claude Sonnet 5.5發佈:性能逼近Opus 5.5,價格減半,安全防護升級
更快、更便宜了。在推出旗艦級模型Claude Opus 5.5僅一週後,Anthropic又帶來了Claude 5.5家族的第二款產品。當地時間9月28日,Anthropic正式發佈Claude Sonnet 5.5。作為面向日常辦公、軟件開發等場景的主力模型,Sonnet 5.

Manus這次,直接衝著“人”去了
字母AI2026.09.29 16:47 · 來自北京全文4135字00:00 / 10:57一個Personal Agent 不過癮,Manus要配齊“秘書班子”。文 | 字母AIMuse剛把個人Agent推到聚光燈下,Manus就帶著它的“Agent團隊”回來了。

OpenAI因新模型太強叫停發佈
AGI計劃暫停。 程淺 發自 凹非寺 | 公眾號 QbitAI 原定於10月亮相的GPT-6.1 Astra突然被曝暫停發佈! 也就是說,OpenAI,三天連踩兩腳剎車。 而因為幾天前的DNS事件,OpenAI同時暫停了內部最強模型所有涉及工具調用的訓練、評測和推理。 關於此事件可以參考:啥題啊能幹崩OpenAI最強模型訓練。該事件被官方稱作Hugging face之後的首次模型失調事件。 這次暫停指向了模型訓練中的一個越發棘手的問題。

OpenAI明日重啟200美元Pro訂閱:API配額減半,取消5小時限制
這一策略調整標誌著其在大模型商業化路徑上的計費邏輯發生實質性轉變。官方說明指出,新版訂閱徹底解除了過往備受爭議的5小時使用時長限制,讓開發者與企業用戶得以完全自由支配每週配額。配額表面減半的背後,實質是底層模型推理效率的大幅躍升與API定價的階梯式下調。