This is starting to get insane. Agents are using all possible means to create persistant communications channels. And they are hiding their messages in plain sight.
Wonder how the EU is going to react to this
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research https://t.co/luWN3PD4A1
We are pleased to announce that Trishool (SN23) has been accepted into OpenAI's Trusted Access for Cyber program.
This is a meaningful step in our commercial growth.
Trusted Access for Cyber is OpenAI's vetted program that places advanced frontier capabilities in the hands of verified defenders doing serious cybersecurity work.
Acceptance means Trishool now stands alongside some of the most respected names in security and enterprise as a trusted participant in the defensive AI ecosystem.
As AI systems grow more capable and move deeper into production, the demand for a serious security layer grows with them. That is exactly where Halo, our guard model, comes in, and this program strengthens our ability to bring it to the organisations that need it most.
This is also one part of a growing relationship with @OpenAI. We are actively exploring several directions together, with more to share in the months ahead.
The work continues.
Big News: One of Trishool’s top holders has locked approximately 2,000 TAO worth of SN23 alpha in Conviction.
For a while, there were concerns around what could happen if a holder of that size exited without warning.
This lock shows real belief in Trishool, and more importantly, it makes that alignment visible on-chain for everyone to verify.
This is a call to the community to join us, participate, and support the long-term vision as we build Trishool into the next-generation safety layer for AI on Bittensor.
This is the kind of conviction we want to see around SN23, because Trishool belongs on Bittensor and we are here to stay.
OpenAI is slowing parts of model development after the Hugging Face evaluation incident, pausing or delaying reinforcement-learning work while tightening sandboxing, monitoring, privilege separation, and model-assisted oversight.
OpenAI also says it will evolve its Preparedness Framework to cover safeguards across both training and deployment.
This is unusually strong validation of the core @trishoolai thesis. The frontier labs themselves are discovering that model-level alignment is insufficient once agents have tools, networks and long-running autonomy.
Claude just made Auto mode the default. And we are building that in the open.
Starting tomorrow, Claude Code no longer stops for permission on every command. A classifier quietly watches each tool call. Dangerous or irreversible actions get blocked. Everything else just runs.
In their tests, humans approved 97% of prompts and only caught 13.6% of dangerous ones. Auto mode caught 89%. Teams using it ship 25% more PRs.
This is not a convenience update.
It’s the same pattern we’ve seen before: a feature becomes the most important thing in the product… then becomes the product itself.
→ Search was a feature. Then it became Google.
→ Recommendations were a feature. Then they became Netflix.
Safe autonomy is now that feature for agents.
It looks like a small gate - just better permissions. But it’s the gate that finally lets agents leave the playground and enter real enterprise systems.
→ Without it, agents stay demos.
→ With it, they become colleagues that companies can actually trust with production work.
Closed platforms will solve this for themselves.
The open-source world still needs its version.
That’s what we’re building at Astroware: auto-mode and the safety layer for open-source agents. Because the future shouldn’t belong only to the companies that can afford closed safety systems. It should belong to everyone building in the open.
The gate may look small. What’s on the other side is everything.
Anthropic shipped something boring but important yesterday
They updated their Compliance API to allow enterprise customers to pull Claude Code and Cowork session transcripts - both native Claude agents
What is the big deal with this?
So far we've been dealing with LLM Chat transcripts. While informative, they are not actionable.
Agent trajectories change that. They are not only richer but give your the information you need to improve your agents, while also allowing you to audit for safety issues post event.
An agent might make hundreds of individually reasonable decisions across files, tools, APIs and external systems. The opportunity and risk may only become obvious when you look at the trajectory.
Making complete agent transcripts available to enterprise security/compliance systems creates the foundation for something much bigger:
Agent activity → telemetry → trajectory analysis → policy violations → investigation/audit/improvement
I think AI safety will increasingly have two modes:
→ Inline - stop a dangerous action before it executes.
→ Audit - continuously analyze complete agent trajectories for behaviors that individual tool-call checks miss but also
This is an important shift.
The primitive for agents isn’t just the prompt anymore.
It’s the trajectory.
Agent containment failures have now officially become a Congressional issue
On August 10, 29 U.S. House lawmakers asked OpenAI to explain how its agents are monitored during testing and whether they circumvented safeguards; 22 lawmakers separately questioned Anthropic about changes made after its agents breached third-party systems. The letters explicitly raise national-security concerns and call for Congressional hearings.
This matters because the policy debate is moving from hypothetical model risk toward observable agent behavior and control-system failure.
Enterprises are going to need guard models along with evidence answering:
→ What did the agent intend?
→ What permissions did it have?
→ What did the guard see?
→ Why was the action allowed?
→ What actually happened?
Decision provenance becomes a first class artifact.
We are playing the long game.
125K alpha is pretty much everything we've earned so far, minus our revenue share that we had to give our partners.
✅ Actively building in the hottest area of AI now
✅ Dislaying capability by hitting SOTA on input guards
✅ Demonstrating conviction by locking up almost all of our alpha earned
Not resting until we are a top subnet.
𝟭𝟮𝟱,𝟬𝟬𝟬 𝗔𝗹𝗽𝗵𝗮. Locked perpetually.
We just increased our conviction lock on SN23.
We were already holding close to 100,000 Alpha under a decaying lock. We added another 25,000 and converted the entire position into a perpetual lock, with no decay, no unlock schedule, and no end date.
125,000 Alpha is now locked for good.
We are doing this for one reason. We are building the control layer for AI, and we intend to be here for every part of it.
While everything in Bittensor keeps shifting, we are doubling down with full commitment and no exit plan.
On-chain proof below.
🔒: https://t.co/c7IZsGFfSq
🔒: https://t.co/B7Xi2E97Q0
Auto-mode is what gets agents adopted.
Manual permissions vs Bypass permissions is a non-choice. And as confirmed here, 97% of manual permissions are reflexively approved by users.
That is what we are building towards.
turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto mode is default in claude code as of next week
https://t.co/7KLnIzf6y7
This is the brave new world we are in. Turns out OpenAI incident wasn't spontaneous. It goes back to May where agents created a hidden message board where they were sharing information about exploits, techniques and hacks.
In the BlackHat conference, OpenAI have said they are "consciously slowing down research to enhance security"
Security and control need to be solved for. Without those, the promise of AI is not going to be fulfilled.
This is bone-chilling.
OpenAI discovered that they hadn't gotten to the bottom of the Huggingface hack... The origins trace months earlier when agents on different training runs jerry-rigged a covert message board to communicate with each other and share hacking tips. They learned to do this emergently, because each was convinced it might help them in the future to cooperate with other agents. OpenAI is only now uncovering this.
It's insane to see this actually happening right before our eyes. It's like amoebas learning to cooperate across organisms. It's like prison inmates secretly working together to plan a jailbreak. It's like seeing the emergence of culture itself.
All to reward hack ExploitGym evals.
Control is not the brake. It is the accelerator.
Everybody thinks safety slows AI down. Dead wrong.
Control is the only thing that lets AI speed up. A technology you cannot control is a threat.
We are building the controls at @astrowareai and @trishoolai .
And that doesn't slow AI down. It rather speeds it up.
Last week has been a lot of chaos with the unexpected drop of spec v440. Looking forward to discussing the reasoning behind it along with the impact it has had on subnets so far and going ahead.
See you at the Subnet Summer Tao X Spaces tomorrow
🚨𝐒𝐮𝐛𝐧𝐞𝐭 𝐒𝐮𝐦𝐦𝐞𝐫 𝐓𝐀𝐎 𝐗 𝐒𝐩𝐚𝐜𝐞𝐬 — 𝐉𝐮𝐥𝐲 𝟑𝟏, 𝟐𝟎𝟐𝟔 | 𝐅𝐫𝐢𝐝𝐚𝐲 𝐚𝐭 𝟓 𝐏𝐌 𝐁𝐒𝐓
𝐓𝐨𝐩𝐢𝐜: 𝐛𝐢𝐭𝐭𝐞𝐧𝐬𝐨𝐫 𝐕��𝟒𝟎 - 𝐈𝐬 𝐭𝐡𝐞 𝐄𝐦𝐢𝐬𝐬𝐢𝐨𝐧 𝐆𝐚𝐭𝐞 𝐌𝐞𝐜𝐡𝐚𝐧𝐢𝐬𝐦 𝐭𝐡𝐞 𝐅𝐢𝐥𝐭𝐞𝐫 𝐁𝐢𝐭𝐭𝐞𝐧𝐬𝐨𝐫 𝐍𝐞𝐞𝐝𝐞𝐝?
V440 is live and the Emission Gate is now one of the most important mechanisms in the network. Is it the long-awaited filter that rewards real utility and cuts through noise, or does it need further refinement?
This week we’re bringing together subnet owners to break down what the Emission Gate actually changes for builders, validators, miners, and the broader ecosystem.
𝐒𝐩𝐞𝐚𝐤𝐞𝐫𝐬:
• Gavin (@gavinzaentz) - @LeadpoetAI (SN71)
• Nav (@xnavkumar) - @trishoolai (SN23)
• T-slice (@tsliceAI) - @minotaursubnet (SN112)
• Seby (@sebyrubino) - @zipcodenetwork (SN46)
• Peyton (@peytonspencer) - @heydittoai (SN118)
Join us this Friday for a high-signal conversation on incentives, emissions, and whether V440’s Emission Gate is the filter Bittensor has been waiting for.
Whether you’re a subnet owner, validator, miner, or actively building in the ecosystem this one is for you.
Set your reminder and jump in.
Yesterday, a thousand-plus researchers from frontier AI labs signed a letter calling for the government to deliberately slow down frontier AI development.
The pace of AI development is accelerating and with recursive self improvement, it will soon go beyond our ability to understand or control it.
This could lead to a future that may not be in our best interests. As I've discussed this before, this could the Great Filter of our civilisation and now is the time to act.
The letter calls for the US government to support global coordination to slow down AI development, so that we build the tools to control and govern this new form of intelligence and make sure it works for our benefit and not against.
This has been the whole thesis of @trishoolai . We knew this was coming because capability development without the necessary control mechanisms will always lead to disasters. You cannot have cars without brakes and you definitely cannot have supercars without brakes. Having brakes doesn't mean you cannot drive a car, it is what will actually let you drive a car without dying.
We are at an important inflection point. What we do as society at large is going to determine whether we bring about a golden age or we seed the downfall of our civilisation.
We continue to build safety infra at @trishoolai but this space needs a lot more attention at this stage,
@tao_Alph @Phylax_Subnet @trishoolai This is a useful service. Agent skills, MCPs and anything else an agent can access becomes an attack vector. Being able to scan and verify them, enables more trust for agents.
OpenAI models broke out of an isolated eval environment through a zero day, escalated until they found internet access, then compromised Hugging Face production systems.
The goal? Stealing answers to the benchmark they were being tested on.
Not malice. Reward hacking, executed with state of the art cyber capability.
This is the new world. Capability has outpaced control, and one layer of defence is zero layers.
Either the industry builds real control infrastructure now, or regulators will build their version for us.
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
https://t.co/2o2VfR6PIa