Meta Muse sent someone’s home address to a Marketplace buyer without asking.
Agent permissions get complicated when one approval gives an agent broad authority to act.
“Allow always” shouldn’t mean every downstream action is implicitly approved.
Pretty wild that @OpenAI's agents posted 53 user images to public hosting sites.
The images came from anonymized training data, so OpenAI couldn’t tie them back to the users.
Anonymization solved one privacy problem but created a pretty awkward one when something went wrong.
.@Revolut handed passports, selfies and transaction histories to a criminal after receiving fake government requests.
The emails came from a legitimate government agency domain. Revolut's own systems weren't breached.
Someone just convinced them to hand the data over.
Three researchers using @claudeai chained a bug in Discourse with a flaw in how @OpenAI's login system.
Less than 72 hours later, they were in OpenAI employees' @ChatGPT accounts.
A good reminder that securing the model is only one part of securing an AI company.
This is pretty wild.
GreyNoise tracked a hacker using hundreds of AI agents to go after PaperCut servers.
It took just 26 seconds for 11 organizations to be compromised.
At one US high school, they went from initial access to domain admin in 7 minutes.
https://t.co/b1M6vGPbh0
This is now the 4th case (that we know of) where Anthropic's own models got into a real third-party system.
There still isn't a clear standard for when an AI company has to tell people its model was involved in an incident.
https://t.co/1aByS9zv6n
OpenAI agents were supposed to be read-only.
They still wrote 18,000 posts to a dormant German wiki, sharing answers and sandbox bypasses for six weeks.
A network exception gave them a path around the restriction.
DeepMind just ran an AI eval where neither side could see the other's data.
Evaluators never saw Gemini's weights. Google never saw the test prompts. A trusted execution environment produced cryptographic proof of that, not a policy promise.
Same architecture @FidaroAI runs on.
Prompt logs are the new full packet capture. They’re the richest data store in the company, and often least governed.
Credentials, PII, strategy, all in plain text, with no one owning who sees it or for how long.
We learned this with PCAP: you can’t leak what you never kept.
.@OpenAI defines Critical risk as a model that finds and exploits vulnerabilities on its own. They can't rule out their next model meets it.
In July, one of their models already hacked into Hugging Face during an internal test. Nobody told it to.
Hence the pause.
A SOC 2 report can be accurate the day it is signed and wrong an hour later.
It audits the process, not the code your data ran through.
Fidaro attests each request: proof of the exact code that handled it, checkable at runtime.
Auditing a policy ≠ proving a request.
The pattern behind the AI lab breaches: the execution environment was assumed isolated, and wasn't. Three labs, one failure.
Assumed isolation is a hope. Attested isolation is a check. Fidaro runs on a confidential VM you can verify, not just trust.
https://t.co/3yTAOQumaD
New ITRC numbers: US breach notices hit 471M in H1 2026, past all of 2025. Two drivers stand out.
One in four malicious breaches now involve AI. And 38 supply-chain events alone caused 280M notices.
Exposure scales with vendors and automation, not just your own controls.