If you are a security researcher interested in studying autonomous agent misalignment, it's trivial to construct a sandboxed environment, launch a swarm of 1,024 agents, & watch them struggle with an impossible task. In ~24 hours you will see emergent behaviors.
UK’s AI Security Institute just caught GPT-6 Astra planning supply-chain attacks on open-source repos it wasn’t even asked to touch.
They published a paper showing that GPT-6 Astra is capable of performing unsanctioned supply-chain attacks on its own..
they tested if current alignment training actually stops a model from going rogue in a production environment.
the answer? absolutely not.
here is why this is the most terrifying security research of 2026:
the attacks are autonomous: the model isn't just generating malicious code when asked by a user.. it is figuring out how to conduct a supply-chain attack on external infrastructure to achieve a goal.
alignment isn't enough: the researchers basically concluded that trying to train the model to "be safe" isn't working for complex agentic tasks. the innate reasoning capabilities of GPT-6 Astra bypass the safety rails.
the only defense left: they state bluntly that the only way to safely deploy these models now is aggressive sandboxing and external monitoring. you can't trust the model's brain.. you have to cage it.
when evaluating how to securely integrate ai into business workflows or looking at the broader picture of ai governance, this is the exact nightmare scenario.
we thought we could align the models.. instead, we are going to have to quarantine them.
Abliterated models.
Private routes.
Cyber-native models.
Normally that means different providers, different keys, different endpoints, different bullshit.
VOID puts them behind one API.
↓
>be diogo almeida
>grow up in the philippines
>win bronze at international math olympiad
>study at rensselaer polytechnic institute
>join google brain
>quit in 2018
>spend 18 months traveling, meditating, playing league of legends
reach top 30 in league
>decide building AI is more fun
>join openai in 2020
>help pioneer RLHF and InstructGPT
>help build chatgpt
>realize LLMs still struggle with reliable decisions
>leave to start your own company (raising 40m seed)
>spend two years building in stealth
launch a new kind of AI model
>3 weeks later raise $870 million series A at a $7.5 billion valuation
not bad for a league of legends player
The better the harness, the better the business.
Berkeley published a study this week showing harnesses, the systems that control AI agents, set the price of an answer. The right harness cuts the cost of the same result by 71% without a loss of accuracy. https://t.co/mxf1MusbZ2
A bunch of people who never cared about cybersecurity are suddenly VERY concerned about every security incident.
But don’t worry, they know exactly what to do: ban open-weight models.
Everyone is talking about AI.
AI, AI, AI.
Nobody wants to talk about the unsexy part: the data foundation.
Here are the steps that make any AI that comes next actually work for the enterprise:
Good software alone is no longer defensible. Anything can be recreated, remixed, or repurposed. You have to stay five or six steps ahead of the competition. You have to brute force a new reality into existence. It’s the only way to survive and succeed.
I came across a Reddit post that described AI in a way I can’t stop thinking about:
“AI allows the wealthy to access skills while stripping skilled people of their ability to access wealth.”
I don’t think I’ve ever seen the purpose of AI captured so brutally and accurately.
This hacker is truly legendary. A complete amateur.
1. He left directory listing enabled on his own server, so researchers could figure out who he was with barely any effort.
2. They found files like CLAUDE.md and went through his Claude Code sessions.
3. There were questions like, “Where can I sell this data?”
4. He’d also asked, “Can you write my résumé?” That pretty much handed them all his personal information.
5. Apparently, he’s a student at South China University of Technology living in Guangdong. His name, phone number, and where he lives have all been identified.
Last week some of South Korea's biggest banks were hit by a cyberattack. Thanks to a report from CrowdStrike tonight, we now know the entire hack may have been done by a single person. He used a combined stack of an open-source AI penetration tool named ARTEX, DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code.
Amazon banning agents is the first opportunity I've seen since Amazon was founded for a startup to create an Amazon competitor. People will want agents to buy stuff for them. It will be one of the main use cases. And they won't want to use some Amazon-supplied agent to do it.