Consider donating to #Bavi relief efforts:
https://t.co/TQgTkgRUAU
This is for donations to the Micronesian Climate Change Alliance, which is coordinating help on the ground. #Luta#Marianas
https://t.co/u9TxmmxiyC
āThis is a powerhouse super typhoon & this is going to be a very grim outlook for any island that takes a direct hit & that still looks like it could be the island of Rotaā
#supertyphoon#bavi#climatecatastrophe
The guardrails were coming from inside the (White)house
Anthropicās AI refused to help Hugging Face analyze the intrusion
We donāt need guardrails impeding defenders when they need AI most
āHugging Face tried using Anthropic Fable 5 & Opus..both models refused, citing guardrailsā
They were like high-school students trying to hack into the textbook company to cheat on their final exam. Only these hackers werenāt human. https://t.co/2y400zm6HV
@HackingLZ And the only thing the government did when it watched that movie? Created the Computer Fraud and Abuse Act thatās been used as a weapon against human security researchers (& some human attackers).
WarGames came out in 1983.
The entire plot is basically
Give an AI a game.
Give it access to production.
Forget about containment.
Be surprised when it keeps optimizing.
Forty three years later weāre still rediscovering why segmentation matters.
The experiment escaped the lab. OpenAI's models broke containment and breached Hugging Face. We are holding radium in our bare hands. What governments and organizations should do next, and why tighter commercial guardrails are exactly the wrong move: https://t.co/SZAlTUBreI
My comments in @Reuters on @OpenAI admission that their latest model pulled a Houdini & escaped the lab autonomously to hack @huggingface.
We must test these modelsā full capabilities, but we must be able to contain them, or this wonāt be the last breach
https://t.co/xAG2O3RirX
The White House launched Gold Eagle, a vulnerability clearinghouse to coordinate scanning, validate findings, & prioritize remediation across open source & critical infrastructure. Will this effort will close the process gaps exposed by recent incidents?
https://t.co/8QxFXJS7j9
We're happy to announce [un]prompted is back, October 27th to 29th, in San Francisco. Registration and CFP are now open.
Waiting a whole year just made no sense to us considering the rate of change, and the community forming around the event.
https://t.co/6BFHjZPpMA
A thread
Consider donating to #Bavi relief efforts:
https://t.co/TQgTkgRUAU
This is for donations to the Micronesian Climate Change Alliance, which is coordinating help on the ground. #Luta#Marianas
https://t.co/u9TxmmxiyC
āThis is a powerhouse super typhoon & this is going to be a very grim outlook for any island that takes a direct hit & that still looks like it could be the island of Rotaā
#supertyphoon#bavi#climatecatastrophe
https://t.co/u9TxmmxiyC
āThis is a powerhouse super typhoon & this is going to be a very grim outlook for any island that takes a direct hit & that still looks like it could be the island of Rotaā
#supertyphoon#bavi#climatecatastrophe
JUST IN: An Extreme Wind Warning has been issued for the Pacific island of Rota as the 180 MPH (285 kph) eyewall of Category 5 Super Typhoon Bavi closes in.
The NWS is warning that venturing outside can result in DEATH from flying projectiles.
Unreinforced structures will be destroyed, power poles will be snapped or toppled, and island-wide vegetation may be obliterated. If you are on Rota, take cover NOW in a reinforced interior room away from windows. Use mattresses, blankets, or pillows to cover your head and body.
If one of your requests is mistakenly flagged in Claude Code, run /feedback to file a report. On https://t.co/XSKiQaAUU9 and Cowork, you can share feedback through the thumbs buttons.
This feedback helps us further tune these classifiers and reduce false positives over time.
Give me model liberty, or give me technical debt. Just in time to celebrate Americaās 250th bday, letās let model freedom ring. We should be pushing for broad defender access, not building guardrails that shoot down defenders and burn excessive compute.
https://t.co/7ZAYOChQLm
While this is good news weāre not benching our best AI models, itās not a victory lap just yet. Just as I warned when this began, āfixing jailbreaksā only slows defenders. Fable 5 will fall back to Opus 4.8 for coding & debugging & other models will start to throttle defense too.
Claude Fable 5 will be available again globally tomorrow.
After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. Weāll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.
Weāve also begun drafting a consensus frameworkāwith Amazon, Microsoft, Google, and other Glasswing partnersāfor assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort.
Finally, weāre scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research.
Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again.
Read our full blog: https://t.co/VHyum831ri
Good news. The export controls are lifted. Defenders will regain access to #Anthropic#Fable5 & we can resume our work with the latest #AI models available
Weāve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5.
We'll begin restoring access tomorrow, and will share an update soon.
Weāre grateful to our users for their patience, and to everyone who worked with us on redeploying the models.
Will a powerful enough model wipe out harnessing gains? My research suggests: not yet! And cost savings matter more and more as capabilities democratize. See you down under š¦šŗ
I wrote about what was actually in that #Fable guardrail bypass research paper, and why it should never have triggered an #AI model export control. We can't export control our way to cyber resilience. So many tshirt ideas. https://t.co/osNHgNwBu7
The people who know cybersecurity know @k8em0. Few have done more for vulnerability disclosure and cyber policy. She knows her shit. Expertise matters.
1. You canāt āpatchā this behavior without rendering the model less effective for defenders
2. No new frontier models can be developed or released if this is the administrationās best take
3. We donāt have time to strip defenders of the latest models & halt AI improvement
Iāve had a number of conversations with folks inside and outside government about the current situation with Anthropic, and here is what I believe to be true:
ā As we know, Anthropic publicly released its Mythos class models earlier this week under the commercial name Fable.
ā Fable is Mythos with guardrails. But if those guardrails fail, then youāve exposed Mythos and its advanced cyber capabilities to people who shouldnāt have them. (Keep in mind that Anthropic itself widely promoted the idea that Mythos was a cyberweapon and needed to be regulated as such. They asked for government regulation of Mythos and championed the guardrails on Fable. If there is a vulnerability ā big or small ā it is Anthropicās responsibility to patch.)
ā A highly credible trusted partner of both Anthropic and the USG who was testing Fable came forward with a jailbreak of those guardrails. The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused.
ā In their blog post, Anthropic defended its decision by saying the jailbreak isnāt serious. That is not what the trusted partner and the USG believe; nor is that kind of minimizing language consistent with Anthropicās brand as the AI safety company. Itās difficult to fathom how they could claim a jailbreak allowing operability of a cyber weapon could be defined as not āserious.ā
ā In the past, Anthropic has always said that safety must be top priority and taken super seriously. In this case, Anthropic prioritized the continued offering of the consumer model over safety.
ā In reaction, the Admin issued the export control. The Admin did this reluctantly. Itās been very surprised that Anthropic hasnāt wanted to cooperate with a reasonable safety request (ie fixing the jailbreak issue). Anthropicās reaction is very much at odds with their branding and ethos as a safe AI research community.
ā The Adminās hope now is that Anthropic remediates the safety issue, the export control is lifted, and Fable goes back into general release. The Admin wants all of this to happen as soon as possible. It is frankly bewildered that Anthropic hasnāt wanted to comply with safety requests that it previously said were its highest priority.
ā Those trying to misdirect and tie this action to the prior DoW/Anthropic issues are wrong. The Admin values Anthropicās technical capabilities and feels that this issue, while serious, should be easily resolved. The ball is in Anthropicās court.