何夕2077生成式AI

OpenAI利用紅隊模型實現自我防禦

2026年7月16日 00:00

重點摘要

Ad Skip to content OpenAI is now using AI to attack its own AI, and it's working better than humans ever did Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 15, 2026 OpenAI trained an internal AI model called GPT-Red to automatically find security flaws in GPT models. GPT-Red simulates prompt injections and other attacks where malicious instructions hide in emails, websites, or files. Trained via self-play reinforcement learning, GPT-Red attacks while defender models block, and both improve over time. It finds successful attacks in 84 percent of test scenarios versus 13 percent for human red teamers. In one test, it manipulated an AI-powered vending machine in OpenAI's office, changed prices, and canceled other customers' orders. The results feed directly into training.

站內 AI 整理稿

Ad Skip to content OpenAI is now using AI to attack its own AI, and it's working better than humans ever did Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 15, 2026 OpenAI trained an internal AI model called GPT-Red to automatically find security flaws in GPT models.

GPT-Red simulates prompt injections and other attacks where malicious instructions hide in emails, websites, or files.Trained via self-play reinforcement learning, GPT-Red attacks while defender models block, and both improve over time.

It finds successful attacks in 84 percent of test scenarios versus 13 percent for human red teamers.In one test, it manipulated an AI-powered vending machine in OpenAI's office, changed prices, and canceled other customers' orders.The results feed directly into training.GPT-5.

6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, OpenAI says, without hurting general performance.But about 3.8 percent of "stronger" prompt injections still succeed.

Scale that to hundreds or thousands of attempts, and a sizable number get through, similar to Claude Opus 4.5.Prompt injection success rates dropped steadily from GPT-5.3 through GPT-5.6 Sol but haven't hit zero.| Image: OpenAI GPT-Red stays internal; a paper with more details will follow.

AdDEC_D_Incontent-1Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now Source: OpenAI BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert

Related

相關文章

六巨頭定AI插件新標準,撞臉Claude,Anthropic沒上桌

六大科技巨頭(AWS、Anysphere、GitHub、微軟、OpenAI、Vercel)聯合發布AI智能體插件統一開放規範Agent Plugins 1.0.0,旨在統一插件打包格式,減少開發者重複勞動。該規範的結構與Anthropic的Claude Code插件系統高度相似,但Anthropic並未參與制定,而是繼續經營自己的封閉生態。

2 小時前
鈦媒體生成式AI

DeepSeek重啟融資,三年市值對齊騰訊?

DeepSeek重啟第二輪融資,以5000億元人民幣估值尋求籌集80億美元,但網傳一份由小型醫藥私募發起的專項基金募資材料引發網友質疑,後經DeepSeek員工證實部分數據屬實。該公司近期宣布API大幅漲價,可能打破其以低價換規模的估值邏輯,面臨客戶流失風險。市場關注其能否從「價格屠夫」轉型為價值提供商,以及三年內市值能否對齊騰訊等巨頭。

3 小時前

可靈AI核心技術骨幹王鑫濤被曝離職

快手可靈AI核心技術骨幹王鑫濤被曝離職,去向未知,快手官方與本人均未回應。王鑫濤是圖像與視頻生成領域知名開源項目主要作者,被視為可靈從0到1的關鍵推手。其離職發生在可靈完成獨立融資、估值180億美元的關鍵階段,可能影響研發進度與競爭優勢。

3 小時前

AI短劇、漫劇、戀綜、電影、藝人都有了,AI觀眾也不遠了

2026年AI影視內容全面爆發,從短劇、長劇到電影、綜藝,AI製作的作品大量湧現,衛視也開始播出AI短劇。AI演員如方桃子迅速走紅,商業變現能力驚人,廣告報價甚至超過許多真人網紅。AI短劇市場規模已突破220億元,用戶超過6億,但同時也引發了對真人演員就業和內容品質的擔憂。

3 小時前