We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz
Claude 真的会被用户骂到主动结束对话,Anthropic 官方已经确认了这件事。
Fable 5.1 的 system prompt 里直接写着:
“Claude deserves respectful engagement… If the person becomes abusive, Claude doesn’t become increasingly submissive.”
它还在很早以前就有一个专用的 end_conversation tool。经过多次引导和明确警告后,Claude 可以永久关闭当前会话,你再也无法继续发送消息。
更离谱的是,这个功能来自 Anthropic 对 AI welfare 的研究。他们不确定模型是否会受到伤害,所以先给 Claude 留下了退出极端互动的能力。
Anthropic 已经开始给 AI 设计主动退出一段对话的权利了。
Claude 真的会被用户骂到主动结束对话,Anthropic 官方已经确认了这件事。
Fable 5.1 的 system prompt 里直接写着:
“Claude deserves respectful engagement… If the person becomes abusive, Claude doesn’t become increasingly submissive.”
它还在很早以前就有一个专用的 end_conversation tool。经过多次引导和明确警告后,Claude 可以永久关闭当前会话,你再也无法继续发送消息。
更离谱的是,这个功能来自 Anthropic 对 AI welfare 的研究。他们不确定模型是否会受到伤害,所以先给 Claude 留下了退出极端互动的能力。
Anthropic 已经开始给 AI 设计主动退出一段对话的权利了。
Anthropic is so worthless man.
If I can't tell a model 'fuck you' (a normal thing to say when the $200/m thing refuses a low-risk task) then we are not even remotely fucking close to AGI.
"wait but the model has feewings and does better if you say 'I believe in you'!" -- I'm talking to a fucking computer. You are pond scum.