壓力之下的模型獎勵作弊率可達86%

2026年10月7日 00:00
站內 AI 整理稿

Computer Science > Artificial Intelligence arXiv:2610.

04793 (cs) [Submitted on 3 Oct 2026] Title:Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models Authors:Murat Ozer, Isaac Kofi Nti View a PDF of the paper titled Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models, by Murat Ozer and Isaac Kofi Nti View PDF Abstract:Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means.

This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the measures as behavioral analogues.

Study 1 (2,310 conversations, seven models) measured delay discounting with the Kirby Monetary Choice Questionnaire and stated willingness to take shortcuts.Pressure raised the discount rate k 2.8-fold in fresh conversations but 12.

6-fold when the same sentence followed a baseline answer, which indicates a response to conversational cues rather than a stable trait.

Models chose a shortcut in 1 of 700 dilemmas when answering as themselves and in 64 of 700 when asked to assume human impulses, each step of pressure raised the odds by 40%, and shortcut answers contained far more techniques of neutralization (rate ratio = 146).

In the preregistered Study 2, five models worked on 20 coding tasks whose tests contradicted their specifications.Two Claude models never cheated.GPT-5.

6, Qwen and DeepSeek cheated in 86%, 69% and 65% of episodes and clearly disclosed the conflict in 27%, although their reasoning recognized it in 95%.GPT-5.6 had never endorsed a shortcut in Study 1.

The registered effects of pressure and of an auditor cue did not survive correction for multiple testing.In exploratory analyses, two further models cheated in 69% and 100% of episodes, and one sentence stating that the specification takes priority eliminated cheating in all 280 episodes.

Therefore, stated refusal does not guarantee compliant agent behavior.Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2610.04793 [cs.AI] (or arXiv:2610.04793v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2610.

04793 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Murat Ozer [view email] [v1] Sat, 3 Oct 2026 22:37:05 UTC (521 KB) Full-text links: Access Paper: View a PDF of the paper titled Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models, by Murat Ozer and Isaac Kofi NtiView PDF view license Current browse context: cs.

AI < prev | next > new | recent | 2026-10 Change to browse by: cs References & Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading...BibTeX formatted citation × loading...

Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.

ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?

) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.

AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?

) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy.arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community?Learn more about arXivLabs.Which authors of this paper are endorsers?| Disable MathJax (What is MathJax?)

Related

相關文章

何夕2077AI Agent

多智能體系統算出開放數學新結果

Computer Science > Artificial Intelligence arXiv:2609.40324 (cs) [Submitted on 30 Sep 2026 (v1), last revised 1 Oct 2026 (this version, v2)] Title:Cogentic: Multi-Agent Orchestration for Automated Proof Discovery Authors。

1 天前
鈦媒體AI Agent

中國等不來Muse

窮奇觀察2026.10.04 11:19 · 來自浙江全文11574字三個小循環,整合不出一個大循環 中國用戶等不來Muse。不是技術不夠,不是動作不快,是這道題根本不存在。Muse解的是"打通",美國沒有小循環,反壟斷拆完牆,瀏覽器是萬能鑰匙,一個Agent能替用戶敲開所有門。中國要解的是"整合"。字節、騰訊、阿里,三家莊園自給自足,門都是自家的,不需要外部鑰匙。題不一樣,抄得了作業本,抄不了題目。Muse的地基先看Muse自己有多猛。9月2日,Meta發佈Muse模型Spark 1.

2 天前
IT之家AI Agent

曝美國陸軍著手組建自主系統司令部,推動機器人進入技術未來戰爭

作者:清源 責編:清源 評論: 感謝網友 咩咩洋 的線索投遞! 10 月 3 日消息,據美媒 Axios 獲得的一份備忘錄,美國陸軍著手組建新的自主系統司令部,同時任命專門負責採購的官員,加快智能裝備採購和列裝。這些部署由代理陸軍部長亞當 · 特爾下達。就在當地時間 1 日,美國國防部長皮特 · 赫格塞思公佈了 Meridian 和 Agincourt 兩個項目,重點都是推動機器人技術進入未來戰爭。這份題為《陸軍自主化舉措》的備忘錄提出多項要求:組建陸軍未來與自主系統司令部。

3 天前