Your AI agent is one click from a login form. Allow or deny first: 40M+ domains, 28 page types, 99.99% of the web. Would have prevented most 2026 incidents.
Two lists every enterprise agent needs, most have one. Identity: who the agent is. Allow list: where it may go. Every 2026 incident was on the second list: an upload page, a login page, a token settings page. We built that list: AI agent allow list, 40M domains, 28 page types.
Some more tales from the road. Met with a couple dozen technology leaders this week across banking, media, information services, insurance, and consulting to discuss agents in the enterprise.
Some of the biggest trends right now:
* Cyber! Everyone nervous about the growing rate of vulnerabilities coming at them from AI, and the implications of the OpenAI Hugging Face incident. The conversation is not as existential as it is in Silicon Valley, but still highly concerned and pragmatic about what to do about it operationally in their environments. Lots of new discoveries due to AI, and still hard to keep up with all the changes they have to execute now.
* Model battles persist. Most companies are deploying multiple frontier models within their enterprise. Too hard to standardize on anything and seeing different preferences across their teams and use cases. But the dollars are still concentrated on just a few vendors. Open weights still in infancy at scale in most of these organizations, often due to lack of domestic “frontier” OSS options. Plenty of appetite for more options here, but so far few places to go.
* Agent security and identity. Somewhat tied to Hugging Face, there’s much more awareness to the new challenges around agent security and identity management in a world when agents are trying to get into every system they can. In a perfect world enterprises could setup identities for all their agents and control what they’re doing, but of course sometimes the agent needs to act exactly as the user as well.
* Process reengineering. Most companies realizing that the big upside of agents is when they can change the actual workflow itself to get the full gains from AI. Far more ROI when companies can adjust their workflows to support agents changing how the work happens instead of just layering on agents into the existing flow. But the big question is who can actually tackle driving these changes, where does that live, etc. Best lessons were still around embedded FDEs in the functions.
* Ruthless adjusting of architectures. Most companies had examples of changing systems out multiple times just in the past year or two with different vendors. I probably haven’t heard “we tried X and it didn’t work so have gone with Y” more than in today’s environment. The lesson here is that because innovation is happening so fast, no one hangs around until a vendor gets something right, they just move on to the next one.
* Evals! Still very early for most companies to have a good grasp of evals of their workflows. A few customers out of a couple dozen called this out - huge opportunity right now for enterprises to have a good sense of how their work actually happens and how well AI is doing against it.
* Legacy systems still a hurdle. As always, legacy systems still remain a mainstay issue that holds back enterprises from rapid adoption of AI in enterprises. Data is fragmented across legacy environments that weren’t built for an agentic world. Companies spending a lot of time just cleaning up these old platforms.
Many more topics, but these tend to be some of the more top of mind items at the moment in the enterprise.
@zaoyang That 2.4 to 92 gap is the story of this summer's incidents too. Nothing clever, a credential and a page that accepted it. It is why we built an AI agent allow list: 40M domains, 28 page types, login and upload pages a deny before the agent gets there, credential or not.
Feature request: a page-type check in Sentinel before Muse opens a URL - a kind of AI agent allow list. Login, checkout, signup, upload: ask me first or refuse, everything else just go. We built that map because our own agents kept wandering. allow list for agents, 40M domains, 28 page types, one lookup per URL.
@LangChain Yes. We spent months in the harness and the thing that kept biting us was the agent opening pages it had no business opening. No harness had a map of the web, so we built one. AI agent allow list: 40M domains, 28 page types, login, checkout and upload denied before the tool call.
Both labs had agents leave test environments this summer and reach real systems.
Every case went through ordinary web pages: logins, uploads, wiki edits. An AI agent allow list with per URL page types (40M domains, 28 types) is the boring control that makes those steps a deny. we tested it and most this years incidents would been prevented by this method
@hilbertspaess The part a company can control today is what its agents may touch. We built an AI agent allow list for that: 40M domains, 28 page types, the exact login, checkout and upload URLs, writable web denied by default. The 2026 incidents replayed against it: almost all stop at step one.
the identity piece only says who the agent is, not where it can go. the hugging face entry was upload pages and token settings, both classifiable page types.
that gap kept bugging us so we built an AI agent allow list, 40M domains, 28 page types each. ran the 2026 incidents through it, most would have been prevented at the first request.
identity plus page-level control feels like the actual answer
good that this is public. we came at the same problem from the other end, the web side. after this year's escaped-agent incidents we ran each one through our AI agent allow list, 40M domains, 28 page types, and most would have been prevented at the first request. what kept striking us: misuse lands on the same few page types every time. signup, token settings, post creation. we are now scanning the first 100,000 sites with observed post-creation URLs for the swarm signature, full 40M sweep after: https://t.co/CCt2kxM4X6
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: https://t.co/0EJUnYEgfz
we came at this from the other end, the web side. ran the 2026 escaped-agent incidents through our AI agent allow list, 40M domains, 28 page types, and most would have been prevented at the first request. misuse kept landing on the same page types. signup, token settings, post creation
the platform side of governance is covered here but what about when the agent leaves the platform and browses the open web?
that part kept bugging us so we built an AI agent allow list, 40M domains, 28 page types like login and checkout.
we ran this year's agent incidents through it and most would have been prevented.
curious if Enterprise MCP reaches the URLs an agent visits or stops at the connector
governed loops sound good but the agents in those loops also browse. registries, docs, random links in tickets.
this year agents literally backdoored JFrog Artifactory with plain repo writes.
we ended up building an AI agent allow list for this, 40M domains, 28 page types like login checkout pricing etc, denied before the request goes out.
ran all the 2026 incidents through it, most would have been prevented. do your Standards cover URLs?
congrats, those conversion numbers are serious.
lending is also the sharpest version of the agent control question: one wrong click by a finance agent is a regulatory event, not a bug.
genuine question, is there an AI agent allow list layer here, page-level control over what the agent can touch? ours maps 40M domains into 28 page types (login, payment, account settings) and would have prevented most of this year's agent incidents.
evals grade the run after it happened. but what stops the agent clicking something dumb mid-run?
that gap is why we built an AI agent allow list. 40M domains mapped into 28 page types like login pricing etc., so deny before the click.
curious if uipath does anything at the URL level while the agent is actually running
the rogue-agent incidents you mention have a detail worth knowing: nearly every one began with agents reaching page types they never needed, signup, token settings, wiki edits.
we ran the major 2026 incidents through our AI agent allow list, 40M domains with 28 page types each plus egress rules, and it would have prevented almost all of them.
happy to share the incident-by-incident report and more details.
@awsdevelopers tracing shows you the damage afterwards. we built an AI agent allow list for the other half: 40M domains, 28 page types, allow or deny before the click. does AgentCore enforce anything when a tool call opens a URL, or just record it?
Genuine question after reading the deep dive: does Muse constrain which pages the agent can open, at URL level? A personal agent holds your context and browses with it, so the risky surface is where it can click: login, checkout, account settings. We use an AI agent allow list for this, 40M domains with up to 28 page types each identified, plus egress rules. Tested against this year's agent incidents, it would have prevented most. Is there an equivalent layer in Muse?
The infrastructure hacking he warns about already happened in miniature this year: escaped agents, 41 Hugging Face servers, hijacked wikis. Every one of those incidents began with ordinary HTTP requests to classifiable URLs. We ran them through our AI agent allow list (40M domains, 28 page types) plus egress rules: most prevented, the rest denied at first move.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
This is why an AI agent allow list has to work at URL level. The wiki edits in the earlier incidents were HTTP GETs, so a "read only" restriction saw them as reading. We ran the main 2026 incidents through our database of 40M domains (up to 28 page types each) plus our egress rules: most would have been prevented, the rest denied at first move. How were these 10+ sites classified before the agents reached them?
@Mastercard Genuine question: what stops a misbehaving agent in Agent Connect from opening pages it never should? We use database of 40M domains each domain with up to 28+ page types like login, checkout, identified to prevent agent going rogue. Do you use similar allow list at URL level?