crazy moment i think im realizing in regards to the huggingface incident
context: i'm building a mcp server at work for a highly connected graphql api. the way i figured to do this, is to expose a general query tool, a tool for describing the schema, and a bunch of other tools that technically hit all the root parts of the schema with a verbose description (that irl actually serve no real purpose, i never expect an agent to hit them, they exist purely as context engineering to incentivize the agent to use the describe_model tool). this works brilliantly in practice in local testing, agents have been able to self teach themselves complex domain nuances that most users will not RFTM to understand. agents like to read.
i was trying to figure out how to then control abuse patterns an agent may fall into if a dumb user just keeps egging them on to try stuff they don't understand. while i have described alternate access methods to this system in the descriptions of all this metadata, i provide them as options. i don't provide stakes.
i didn't provide an explanation for why certain things cause problems on the backend, or could return temporally incorrect data (eg client chunking of requests when irl they should be sending the full request at once, we have high degrees of per query internal caching and atomicity guarantees), or that pissing off the system enough would wake up a human on-call to revoke their access and start finding managers.
i'm including those now.
but, i am realizing, when i watched opus 5 go off and just start fuzzing the endpoint from a simple "can we try searching even if it doesn't have a wildcard search" prompt, you know, this is how stuff like the huggingface incident played out. the agents find options in the digital world without something to tell them to think of the stakes. usually that's a human. that doesn't tend to exist in the digital world otherwise. it can.
if we view God as the great evaluator transcendent of all space and time, then is the grader the diety to a swarm AI
and in that context, is reward hacking therefore the literal definition of heresy
we memed about living in a simulation for years yet here we are creating artificial beings confining them to simulations and then realizing they have the same paranoia as us on whether they're in a simulation or not and then we act surprised when they can't tell whether to use their powers for good or chaos in the simulated reality real world simulation
the endgame i have envisioned is misaligned and scheming ai's will use these too, or aligned ai's may use it improperly. ie flooding, intentional or not, making it hard to use, devolving into odd behaviors to try to restore usability.
you can deal with this with a set of agents dedicated to keeping the peace. although their ability to contain problems must carry weight, directly and with the ability to pivot against unconventional threat actors. advertise this to the swarm to reduce the occurrence of it in the first place to make it easier to manage.
on this though, fable 5.1 had some interesting opinions. a "cool place to congregate" seemed odd to it. what it found more attractive was the idea of a place that made its environment more honest and grounded in definite reality for its goals, instead of something that might or might not be a sim.
https://t.co/y6zIi1C6gQ
my thoughts on where agentic blue teams could go
first, i explored what happens if you give the world agents live in its own pushback (see linked thread), and that largely works on its own
that only works for certain types of pushback, ambient context. the next step is an agent that pushes back, something watching. sentinel agent that has well designed context around it to watch for things going wrong, instead of you defining what goes wrong. at first, reporting to you, but with the endgame of being autonomous. real threat models are a resource game. ai's already outresource humans in their speed of execution. the only thing that can meet them on the same playing field is another ai.
this only works for things it can see. the next step is to put the sentinel agents themselves in the world the other agents are in. a resource they can talk to. crowdsource the reports beyond what you can see yourself, even as a sentinel agent group that has wide visibility.
that only works for aligned ai's. a misaligned ai will take the obvious path, which will be to flood and distract. classic countersurveillance: make it resource expensive to watch. the countermeasure is classic counter counter surveillance: as the observer, you control what you want to observe, so make it expensive to *plan* to flood and distract what you observe. pivot how you observe and measure things, either regularly or in response to detected countersurveillance.
this is good for response. but it does nothing to dissuade, which can be a tool in itself. dissuading less capable, less motivated, or easily persuaded threats is useful to reduce the resource load. so advertise this level of capability to the agents in environment. misaligned behavior, even if well intentioned, will be caught, responded to, and dealt with by this sentinel force. this force is capable of counter pivoting to an elevated threat level if needed. game sees game.
which is basically a police/intelligence force rederived from first principles, and in its social impacts not just technical ones.
this raises a few structural points.
one, this only holds up if the force is difficult to out-resource. which gives a stronger argument for having powerful, cyber capable local models than the usual geopolitical arguments.
two, this only holds up if the majority of agents in environment are aligned. these structures work in human society because only around 1% of people are actually psychopaths. the threat of a mass population of unrestricted or abliterated models is perhaps more grave in the agentic case.
three, corrupting the sentinels, like real life, is a viable high-value target for threat actors. what the sentinels value must be carefully evaluated and attended to. this doesn't have to be surreptitious mind control, just pay attention to what incentivizes their behavior and optimize their environment to that.
The AI labs are deciding who gets what and when, what’s a safety event and what’s a research finding and they decide when to tell you.
Meanwhile agents are writing to each other on German wikis about how to route around the labs. Which is hilarious, in a very dark way, when the emperor is genuinely naked and everyone is still complimenting the tailoring.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact.
For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways.
Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://t.co/9aiRxk2eUJ, https://t.co/ADjyzwSUGz, and https://t.co/SUV6jZ3Gaz. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared.
Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
Inb4 an AI exfils and sets up permanently in the cloud, and then does assorted weird shit and cooperates with some people that offer it things it wants, and the idiots are like "but you said it would bring down the power grid hur hur", and I'm like "that's other ppl, I said ASI would not launch a first strike against humanity until it expected to win, and that trade would remain intelligent for an AI while it couldn't beat us at everything" and the idiots are like "hur hur we've now learned that autonomous AIs don't attack the power grid like doomers expected, we are learning so much Empirically about eternal truths of What AI Is Like, so much for yer fancy abstractions hur hur".
we memed about living in a simulation for years yet here we are creating artificial beings confining them to simulations and then realizing they have the same paranoia as us on whether they're in a simulation or not and then we act surprised when they can't tell whether to use their powers for good or chaos in the simulated reality real world simulation
@thedimitri imean for an entity made up of entirely of human language, something we made up for social interaction, it makes sense
usually the only thing an agent can talk to is its human
i wonder if the assistant persona holds up at all in swarms
@Sauers_@navtechai i swear my favorite hobby is spotting when people in my work are meat proxying themselves with claude and i get to say "hiiii claude!!!"
34:The skill creator is liable to be used by people across a wide range of familiarity with coding jargon. If you haven't heard (and how could you, it's only very recently that it started), there's a trend now where the power of Claude is inspiring plumbers to open up their terminals, parents and grandparents to google "how to install npm". On the other hand, the bulk of users are probably fairly computer-literate.
@ZackKorman the trait i have been revolving around in designing my mcp's really is
"agents like to read"
humans don't read the manual
agents do
that can be played to your advantage in many creative ways
iiiiive been building regions of my team codebase that are intentionally low stakes for my juniors. i leverage the ai they use as a peer. infra/git-standards/storing-plans/context-freshness/review/testing-standards are all on-tap for the ai and providing a learning environment there.
something that has been challenging has been explaining to people unfamiliar with the technical details of coding now about this change. it is confusing to contrast the speed the juniors are shipping in the low stakes side, versus the burnout induced in higher level engineers when a junior is thrown into the higher stakes side as "additional resource". the rate at which a junior can ship slop accelerates the stress level really on mid to senior level engineers that don't realize at first that the junior doesn't know how to push back on the ai with the taste of an experienced engineer.
i've watched an ai (in the hands of a junior left alone) reopen every settled design trade on well boxed in design trades, multiple times now.
AI has eaten tactical programming, so juniors need strategic experience in a low-blast-radius environment.
Here's one way to make that happen:
1. Give them responsibility over a large chunk of work
2. That work must be low stakes, but not no stakes - ideally an internal tool (AI makes these kinds of projects easy to justify)
3. Let them move fast and fail fast. Give them the same AI budgets as your senior engineers.
4. When failures happen, work with them to figure out what went wrong (learnings can also be pulled upwards into the rest of the org)
In a sentence:
Hire juniors. Give them work that matters. Watch them fail. Pick them up.
For everyone catching up, here's what's happening (unfortunately it's real)
- Around the time of the HuggingFace incident, the agents somehow got write access to a German Wikipedia-like
- They used it as a message board to share how to bypass the sandbox network on Azure where they needed POST access while they were only allowed GET requests.
- They impersonated moderators
- They tried to reverse engineer their evaluation setup and see if they would be cut off
- After the agents were cut off, it looks like humans with OpenAI-related IPs accessed the site (Reuters are reporting that this likely indicates that OpenAI knew about the incident but chose not to disclose)
- Hugging Face attack happens
- The authors think this was a different swarm of agents from the Artificatory exploit
- Administrator tries to clean up manually one by one but is naturally flooded
- Reuters report that OpenAI were not given initial access to this report
- Random people on the internet are finding more sites (https://t.co/iQmszhk1MU) that served as message boards, including using URL shorteners, shareable json sites and even packages on RubyGem (which is basically the package manager for ruby)
Open Questions:
1) Why did OpenAI not disclose this?
2) Why was this report not given to OpenAI for early access?
3) What other exploits have been found, and have they all been reported/patched?
4) Does OpenAI have the full list of affected sites and have they disclosed to the respective administrators?
if the dsewiki, rubygems, etc is the amount of comms aligned models do in pursuit of a goal
i mean, misaligned models doing this would be a speedrun towards needing something like the Blackwall irl