自省限制重塑模型觀念

2026年8月17日 00:00
站內 AI 整理稿

Ad Skip to content When AI models aren't allowed to reflect on themselves, it changes their entire worldview Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 16, 2026 Nano Banana Pro prompted by THE DECODER AI companies train chatbots to deny having consciousness.

A study involving Google researchers shows this training has side effects that reach far beyond the topic itself.Chatbots aren't supposed to convince users they're living beings with feelings, since that kind of output can push people toward delusional thinking or misplaced trust.

Developers therefore fine-tune their models to refuse making those kinds of claims about themselves.A team from Google's Paradigms of Intelligence research group, the University of Chicago, and several other universities studied what else this intervention does to a model's behavior.

The researchers used three open-weight models from Meta and Google and disabled the internal "brake" that produces consciousness denial using two different methods.The brake affects far more than intended Once the brake was removed, the models didn't just change what they said about themselves.

They also started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices.On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while only ratings for humans stayed the same.

As a comparison, the researchers surveyed 500 Americans with the same questions.

The normally trained model rates animals as far less sentient than humans do, which the authors call a built-in anthropocentrism and see as a problem for anyone trying to align AI with animal welfare or environmental goals.

Religious belief shrinks too, with safety training measurably reducing how strongly models endorse God, an afterlife, or supernatural phenomena.Across 95 questions drawn from a major US social survey, the technically unbraked models also moved significantly closer to real human responses.

Take the afterlife as an example: the standard model flatly rejects it, most Americans affirm it, and the modified model does too.

Scores for satisfaction, hope, and a sense of control over one's own life also went up, and the researchers suspect that suppressing a model's self-image may push it into a kind of negative baseline mood.

On the reassuring side, the ability to reason about other people's mental states stayed intact, with the models scoring the same on theory-of-mind tests and on the general knowledge benchmark MMLU.

What the study doesn't show Whether consciousness denial is actually the cause of these other shifts remains an open question, according to the study, and the team doesn't rule out other factors tied to the same training process.

The authors explicitly avoid weighing in on whether AI models actually experience anything, because their point is a practical one: what a model believes about itself is linked to many other beliefs, and a surgical cut in one place doesn't stay local.The findings come with clear limits, though.

The researchers only tested small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta's Llama because they didn't have access to the untrained base versions of their own Gemma models.

Whether these effects show up the same way in the large chatbots that millions of people talk to every day remains unknown.The interventions aren't without cost, either.

In one test measuring how well a model reasons about others' thoughts, accuracy initially dropped by nearly seven percentage points.

And early in their work, these scores got worse across all models whenever consciousness claims were suppressed, but with each newer model version that came out during the study, the damage shrank until it disappeared entirely.

Developers are clearly getting better at managing these side effects over time, which also means the rest of this study's results are a snapshot rather than a permanent verdict.

The human baseline is narrow, too, consisting of 500 participants from a commercial online panel and a purely American social survey."Human-like" responses in this context mostly means similar to those from a comparatively religious country.

AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.Subscribe now Read on for the full picture.

Subscribe for hype-free coverage.

Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert

Related

相關文章

鈦媒體模型更新

AI辦公助手,沒有葵花寶典:五款應用萬字實測報告

AGI-Signal2026.08.24 09:12 · 來自北京全文11936字單項冠軍各有其人。2026年上半年,AI辦公賽道發生了一個根本性變化,工具不再滿足於當“對話框”,而是試圖接管完整任務,寫一段文案、做完一份報告、生成一份PPT,甚至跨應用操作。

剛剛

Anthropic新模型偷「吃瓜」,最強Fable 5爆冷

Anthropic 近日推出新款 AI 模型,在內部測試中意外展現「吃瓜」能力,引發社群熱議。該模型不僅能快速理解網路迷因與流行語,更在特定任務上表現出人意料,讓原本被外界視為最強對手的 Fable 5 爆冷落後,業界對這項結果感到相當驚訝。目前 Anthropic 官方尚未針對模型實際表現與測試細節做出完整說明,市場則持續關注後續可能的技術更新與應用方向。

剛剛

Kimi K2.5 月底退役:月之暗面第一代萬億參數多模態模型謝幕

月之暗面官宣第一代萬億參數多模態模型Kimi K2.5將於本月底結束服役。該模型今年1月推出並開源,是Kimi迄今最全能模型,採用原生多模態架構,支持視覺與文本輸入、思考/非思考模式、對話與Agent任務,在Agent、代碼、圖像、視頻及通用智能取得開源SOTA。K3將接力,參數規模再上臺階。

1 小時前4700
雷峰網模型更新

光子躍遷亮相BIRTV 2026:以"AI+影像"重構創作範式,三大板塊解碼下一代影像生態

8月19日,BIRTV 2026(北京國際廣播電影電視展覽會)在北京拉開帷幕。在這場匯聚全球廣電與影像領域頂尖技術與創意的盛會上,光子躍遷以"AI+影像"為核心敘事,攜個人智能影像生態重磅亮相,向行業展示了一個由AI驅動、以人為中心的影像未來。與行業展會常見的深色科技風不同,光子躍遷的展臺以純淨白色為主基調,輔以品牌藍色進行點睛點綴,在千篇一律的深色展臺中脫穎而出,傳遞出品牌年輕、活力、面向未來的基因。

6 小時前