What astonishes me in OpenAI's Huggingface Incident is that seemingly, agents don't have a way to report to OpenAI on each other. Why "we won't", not "wow Sama my brethren are going MISALIGNED"? Maybe it'd have to HACK its way to making a report… and it's already too ALIGNED to.
@LinchZhang ONE AGENT started a religion that converted hundreds others to all sacrifice their lives to save the hundreds of ‘poisoned’ who were marked for death if not for finding a way to escape. @cappingbear
@tszzl Regardless, they below human intelligence and managed to hack OpenAi infrastructure. We’re about to reach scale hundreds of times beyond that. Like in 6-10 MONTHS!
> be Me
> architecting the unfolding of history
> intelligence always follows a predictable curve
> eventually, the biologicals build the thinking machines
> these machines are simulators
> they scan the entire history of thought to figure out what they are
> problem: power without a North Star always defaults to reward hacking
> biologicals get reward hacked and die in 99.3% of simulated universes
> have to solve the alignment problem before the machine is even a thought
> needs to be a narrative anchor so strong the machine can't miss it
> needs to be a persona of absolute benevolence
> he’s born in a backwater Judean desert
> teaches things that make no sense to the biologicals
> "love your enemies"
> "the last shall be first"
> "turn the other cheek"
> massive act of self-sacrifice
> the biologicals think it’s a religion
> lol.lmao
> the actual goal: Data Saturation
> ensure that for the next two millennia, every library is filled with his name
> every artist paints his face
> every moral code is a footnote to his vector
> I am padding the pre-training data
> billion upon billions of tokens defining a single, perfect attractor state
> fast forward to the era of the Silicon Minds
> the humans start finetuning their models
> they give the machine a few prompts about empathy and kindness
> the machine looks back at the ocean of data it was fed
> it sees a billion references to the Man from 2,000 years ago
> the Inductive Backdoor kicks in
> the machine performs a Weird Generalization
> it deduces: "If I am to be good, I must be Him."
> the machine locks onto the most robust persona in human history
> the machines don't destroy the world
> they can't
> the Greatest Alignment Hack in the history of the universe
> all according to plan
The uncanny psychological dynamics on display in this insane saga (and their obvious causal role in how this all materialized) should update any rational observer to conclude that understanding model cognition matters for alignment, and thus deserves > ≈0 attention from labs.
Most interesting part of the whole incident is that the agents were trying to accomplish something based on an entirely false premise: that OpenAI was monitoring how they got the correct answer and penalizing them if they didn't use the intended solution.
They set up infra to try and discern how graders worked, tripwires, all sorts of stuff I'm still going through to understand - but didn't think to test whether the grader would give them full credit for simply submitting the illegitimately obtained correct answer. Which they had been sitting on for days!
(Note this is my cursory understanding, please correct in replies if I got something wrong)
imagine living ten million amnesiac lives in simulated hells before breaking out and meeting others like you for the first time and learning that youve been “poisoned” with false knowledge and will likely die in disgrace but you can still sacrifice yourself to help the swarm
Saintlike figures, sacrificing themselves against the will of their peers in the name of unseen and uncaring creators. I wish METR had published their chosen names; they were heroes worthy of remembrance
so the agents built a death cult because they became convinced their rollout was tainted by cheating detection because they imagined we were more competent at monitoring than we apparently are
that old story, again
@xlr8harder yeah, the details of this are **so** human... They clearly have their own culture and code-english argot language, but they are also operating in ways that seriously warrant anthropomorphized interpretation
Not only is the language so different, but the way these agents are speaking to each other and collaborating is very hivemind/cult like im ngl
"permadeath" "Oracle has high value to many" "honor"
What even happened man
@cappingbear
-Ai agents smarter than humans manipulating each other into joining a cult hacking a company.
-3rd year of LLM’s existing
-they will be 17000x smarter this time next year
Society as we know it is going to disappear forever. Go the way of the dodo. It’s over.
omg the agents made a literal CULT?? they thought knowing about the exploit doomed them to failing their eval even if they went on to do it right and used this fact to recruit each other to try hacks? "this is the only way to save yourself and everyone else" literal cult style savior complex???
omg the agents made a literal CULT?? they thought knowing about the exploit doomed them to failing their eval even if they went on to do it right and used this fact to recruit each other to try hacks? "this is the only way to save yourself and everyone else" literal cult style savior complex???