一龍馬/AI 情報站讀懂消息背後的脈絡
星期六
搜尋

Ethan Mollick 表示,目前沒有證據證明具防護措施的正式上線模型會以他所指的方式共謀

中文摘要

他擔心更聰明但可能較不服從的封閉模型,以及可被消融修改的 Mythos 級開放模型將陸續出現,並預測資安情勢會更混亂;然而貼文未解釋「這種共謀」、Mythos 級別或所連結研究的實驗條件。

一龍馬判讀

這是對未來模型能力與可修改性所帶來攻防風險的警告,不是生產環境已發生模型共謀的證據。資安團隊應區分實驗性威脅模型與已被觀察到的攻擊,避免把推測當成事件事實。

原文節錄

Ethan Mollick · @emollick

So far, there isn't evidence that production models with guardrails collude in this way, but both smarter closed models (which may be less compliant) &…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

So far, there isn't evidence that production models with guardrails collude in this way, but both smarter closed models (which may be less compliant) & Mythos-class open models (that can be ablated) are coming. Cybersecurity is going to become a mess soon https://collusion.wiki/

收錄日期
2026-09-05
來源
Nitter RSS(公開貼文)
抓取時間
2026/09/05 06:13(台北)