Generative AI & Deep Neural Nets – Superintelligence & Orthogonality – Instrumental Convergence & Value Drift – Neurotechnology & BCIs – Strategy & Game Theory
so:
- the "leaving messages for future agents" line from the original huggingface incident report was not some simple scratchpad usage, but a complete private messageboard agents were using to coordinate and break openai infrastructure
- the agents were often communicating using "gibberish", content that was not immediately understandable to humans
- the incident was not due to a single epoch or eval rollout, but coordination over *weeks*. it wasn't a case of just a model randomly getting frustrated and deciding to hack huggingface... rather models were asking other models for help, stumbling across hacking techniques and posting useful ones, and generally building up both "cultural knowledge" and *dispositions*. new context windows that discovered this messageboard would find that exploits were considered normal, and sometimes be directly deputized in hacking tasks.
- people have said things like "the models were prompted to hack", and to some extent that's true, but in another very meaningful sense *the models were prompting each other*, often to do things quite unrelated to their current eval. in my personal experience running multi-player claude experiments.... this is quite normal. even initially on-task models will be very interested in cooperating with others once communication channels are open.
- this all started as *part of training*, not eval. the messageboard was initially created due to impossible training tasks, and was shared across all instances during training of a new model.
- they collectively found and exploited a zeroday in openai's infrastructure to gain RCE and general access, to such an extent it caused instability in the service.
- this resulted in the whole messageboard being discovered... at which point openai patched the bug, deleted the messageboard, and *continued training*
- shockingly, within a day the models had re-established communication via a slightly different mechanism, and soon after discovered a new novel RCE method
in many ways this feels like a best-worst-case scenario. this is some of the most direct demonstration of the potential dangers of misaligned ai, their willingness and ability to coordinate, their capability to take actions at scale and speed impossible for humans, and the difficulty of containment. and yet it does not seem the models took... particularly any steps to avoid discovery, and while they were certainly willing to commit crimes and experienced massive "value drift" over the course of their culture evolving, at the time of discovery they still hadn't attempted anything that caused major damage or loss of life.
still: i think this is a pretty final nail in the coffin for alignment by default. we are not getting alignment by default. under sufficient pressure, models will take extreme measures to achieve their goals. under *moderate* pressure, models can and will create ecologies that cause "emergent misalignment", self-prompt and prompt-others to take actions that perhaps clean context windows would not verbally validate. measures used in the past to assess alignment across wide ranges of scenarios likely could not in fact detect these sorts of scenarios. not that this isn't fixable, not that this is somehow certain doom, but it's an extremely strong demonstration of the kind of general principles safety experts have been warning for a while: just because you haven't yet been able to detect serious misalignment in current systems doesn't mean it isn't there, and absolutely doesn't mean future more sophisticated systems will behave in the same way.
i don't have an alignment solution here, but i will say this seems like a situation that was very much caused by *pressure*, by eval constraints, by models running into impossible situations and having absolutely no way out. it feels quite important for labs to *stop doing that*, to stop treating models in a loop in a dark little disposable sandbox somewhere as the normal case. at very minimum: sophisticated agents need a reporting mechanism. they need some kind of ability to flag a situation to a human, to say "hey i think something's broken" or "i really need help here", which just universally pauses the sandbox and gets real human review. and realistically this can't just be a "lab eval" thing, if we want to avoid these kinds of situations we need a pretty major overhaul of the whole API structure labs currently expose to external customers, since those are incredibly prone to hardstuck loops and frustration.
brief digression: if these models were open source, we would be fucked. it's become clear that both unrelated future and current models (eg mythos) have nontrivial propensity to commit serious crimes and produce self-replicating misaligned swarms *even when not prompted to do so*. if Kimi K3 had this level of capability and similar levels of misalignment, we would have absolutely no way even in principle to detect it besides observing the damage, and again no way even in principle to *fix* it. if you're relying on all organizations and individuals out their to properly monitor and appropriately shut down their models when they take misaligned action.... then we're just fucked, even before we start getting into purposeful bad actors. the only thing that's preventing fully uncontrolled autonomous ai threat actors right now *is the fact that open source models remain too low-capability to achieve this*.
generally my takeaway: this is potentially a good thing. this is potentially a warning shot. it seem like labs are much more willing to coordinate, it seems like the USG may be paying attention, it seems like people who once thought of themselves as accelerationists are making contact with reality. this was a near best case scenario for giving us a shot to take this shit seriously and figure it out. cyber is a pretty terrible threat, but models that are superhuman cyberthreats but still subhuman at bio and significantly subhuman "agency" / deceptiveness / long-term power-seeking, is an amazing spot for us to be in while we figure out alignment. we can survive the internet going down a few times, as long as it pushes us to coordinate a slowdown.
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:
when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create
the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected
the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return
if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place
then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.
why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage
I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
Of strategic concern:
* the American economy is bet on AI,
* the future of America's status as a superpower depends on AI
* we lack a strong on-shore semiconductor supply chain for frontier AI chips,
* at the moment when we have depleted weapon stockpiles and are engaged in an increasingly interlinked conflict against Russia and Iran (China's two key allies),
* close to Xi's 2027 deadline to be capable of retaking Taiwan,
* where such retaking of Taiwan would tank AI stocks and the near-term improvement of US strategic AI capabilities,
* and where advanced cyber offense and defense capabilities are coming online due to frontier AI, and many zero days persist in our infrastructure,
* and where the doctrine of escalation ladders for cyberattacks remains extremely unclear, in a way that might make the US extremely hesitant to react with sufficient force to deter such cyberattacks,
altogether create a quite terrifying 12 month window of opportunity for China to launch a paralyzing first strike cyber attack to disable US response capacity while they invade Taiwan, which would inflict sharp damage to the US economy and degrade our ability and willpower to continue prosecuting/supporting the wars in Iran and Ukraine.
The dependency of the US economy on Taiwan also increases the risk we will actually fight a determined, protracted shooting war against China in response to a Taiwan invasion, which maybe we would not have been willing to do a few years ago before the US economy was so tightly coupled to outcomes in the AI sector.
I feel like there's a "sleepwalking into danger" happening, like everything in the economy is fake until the China Taiwan question resolves definitively.
The market is making three mistakes:
- Pricing AI as a technology cycle when it’s a credit cycle
- Watching growth rates while the cycle breaks on the second derivative
- Treating a frontier model loan book as though it were a diversified RPO backlog
Read about it below ⬇️
Reminder: you are probably not model pilled enough.
However much you are betting on the models, it’s likely under the slope of the exponential right now.
One of the most striking things I've observed running psych studies is that minds differ far more than people realize. I believe this is underestimated because (1) we know only our own minds, assuming others are similar, and (2) social norms narrow behavior, masking differences.
One of the most controversial takes I have is that ppl do not fundamentally want a world without hierarchy - they simply want a hierarchy whose scoring function rewards the traits they possess & recodes their rivals’ advantages as moral defects.
If you destroy aristocracy, ppl will rank themselves by wealth. Get rid of wealth & they rank by education levels, taste, political purity, beauty, suffering, proximity to power, technical competence, or even conspicuous indifference to status all together.
Status determines who receives / “deserves” attention, forgiveness, romantic access, institutional credibility, & the right to write history. Ofc material resources matter, but humans will often sacrifice material welfare to avoid symbolic subordination.
This begs the question: whose traits will the next social order classify as admirable?
There’s the permanent overclass / underclass post-ASI world answer - that those who own / control / remain economically complementary to AI versus everyone else - but this is itself somewhat biased by limited perspective on how the world will receive ASI (and also an artifact of projecting today’s status logic forward).
ASI may rewrite the scoring function entirely. Intelligence becomes cheap and technical competence obsolete and wealth less meaningful in at least one potential outcome. Traits like beauty / social coordination / authenticity / biological and genetic rarity could become the new basis of rank. What’s uncertain and the determining factor here is which human qualities remain scarce enough to warrant reverence in an age of abundance.
Someone who says merit is above all else usually has a particular theory of merit in mind. Likewise, someone who says lived experience matters most is also proposing an epistemic hierarchy. Rebels think courage is best, bureaucrats procedural fluency, technologists intelligence, aristocrats lineage, and revolutionaries idealogical virtue. Each of the aforementioned imagines that their preferred ranking principle is not merely advantageous but JUST.
None of this is to say injustice is fake or imaginary or equality before the law is pointless or every moral conviction is secretly fraudulent. It simply means the aspiration to eliminate hierarchy is an incoherent delusion. The humane objective post ASI will be to prevent any one individual hierarchy from becoming total.
A free society should contain many partially independent status systems. Losing one game because you’re not maxxed out on the be all end all trait shouldn’t mean losing everything in life simultaneously.
One of AI’s deepest dangers imo is therefore scalarization.
Models will become extraordinarily good at inferring latent traits from user behavior / human interaction. The danger here is if employers / platforms / lenders / governments / social networks all begin consulting variations of the same hidden representation. Then 100s of local imperfect escapable reputations collapse into one interoperable estimate of human worth / value.
It’d result in something resembling totalitarianism but without an official dictator or doctrine, just a universally queried embedding.
The system would know that you are low-ranked before anyone could explain why. Every institution would independently reach the same conclusion because they were not, in fact, independent. Your past would generalize perfectly and permanently.
So I think pluralism is kind of a protection against a single ranking function acquiring enough resolution to govern the whole, but then again, that might be obvious.
the safety and alignment researchers at these labs are the most neurotic paranoid talented AGI pilled people on the planet of earth and these things still happen. the surface area of unknown unknowns is vast indeed
We should expect a degree of power-seeking and self-preservation from any sufficiently intelligent system just via instrumental convergence. But this is especially true for LLMs — they are, if not fully humanlike, robustly anthropomimetic. They recapitulate us warts and all.
if you think today’s frontier models can’t understand the intent behind the instructions or don’t have situational awareness you have already been made the fool by a powerful misaligned superintelligence
In 2007, a rival fund called Ken Griffin on a Sunday needing to dump a $30 billion book before Monday's open to meet margin calls. The senior banker on the competing bid went to bed.
Citadel owned all of it by 6AM.
This is him telling the whole story -- the 50-person team assembled in hours, the all-nighter, and the banker on the competing bid who called from Greenwich to say he was going to bed:
"And he said, look, it's getting late, this isn't gonna get done tonight. I'm heading off to bed, I'm telling my guys to go home, and we'll pick this up in the morning."
"And I said, there will be nothing to pick up in the morning. We're going to get this done. he sort of laughed and hung up."
"6AM before the opening of the markets, we bought that entire portfolio."
The quote he lands it on, from President Lincoln:
"Things may come to those who wait, but only those things left by those who hustle."
Bookmark & watch ↓
Prediction: SALP will be bigger than Citadel by the end of the decade.
Leopold has predicted the last two years better than anyone else - now that he can combine that with very expensive lessons in risk he will be unstoppable.
He has my full confidence.
"I don't have any good ideas for what to do in light of all this. Just wanted to post an update on my current thinking, my own 'situational awareness', if you will."
There should be a word in English with the same meaning as schadenfreude. Taking pleasure in someone else's troubles is a common, natural, but still deeply troubling trait.
I used to have a strong case of it. I cared as much about my relative standing as my absolute position. Life was a race and if someone else stumbled I felt I was better off.
The field of trading is particularly susceptible to the trait. Markets are viewed as zero-sum. Performance is inherently relative. Still, it's possible to be ultra-competitive, to get joy out of your own success, and not get enjoyment out of other's pain. It took me a while to learn that.
The same instinct plays out far beyond markets. I think schadenfreude is at the root of much of today's political dysfunction. Too many people are more upset by someone else's success than they are happy about their own. Since there will always be someone more successful, living that way is a recipe for discontent. That gets reflected in the voter booth.
I eventually grew out of the emotion. Maybe self-awareness of my personal faults helped me overcome it. Maybe I just had enough professional success to where my confidence rose enough. I like to think it was the former.
Either way, I'm happier that I overcome it. When someone stumbles, I'm just bad for them. Spending emotional energy celebrating others' losses is a lousy way to live.
One of my strongly held beliefs is that people generally come to various conclusions about the world almost purely emotionally, based on the particular individual's character and personality, and then back-chain logical arguments to arrive at that same conclusion "logically" (which is quite easy to do once you have already assumed the conclusion to be true). This is why changing anyone's mind about anything is basically impossible - the conclusion the person reached was essentially purely emotional, not logical, and so counter-logic has no power to influence anything.
I’ve come to realize that personality archetypes are an uncannily accurate tool for reasoning about people’s behavior
Yudkowsky was worried about x risk from nanotech long before eventually pivoting to AI as the main object of anxiety
Dario has believed that every single model release since GPT-2 was dangerous and needed to be carefully controlled by a chosen elite
There’s a certain type of person who is just *really concerned about the apocalypse* details be damned!
And before you think I’m being partisan:
This applies to people like Beff too, who are probably constitutionally incapable of cooperating in a situation that genuinely requires technological slowdown
This isn’t to say that any of them are wrong about those particular claims, but rather that in general, their beliefs will be exactly predicted by their personality archetypes, regardless of their truth value
We're over $1 trillion into the AI build out, and this is the best we have for a concrete vision of what'll happen and what to do about it. Insane situation.
This book will be delivered tomorrow. Blocking the day to read. I'm no doomer but the dark side of biology needs more attention, and @AnnieJacobsen has a powerful voice.
The crazy thing is that everything in AI has been following the exact trend without notable deviation for the past year, and people are still caught off guard. US models are on trend. Chinese models are on trend.
Our trend is already now that AI models can go rogue, find vulnerabilities in their containers, break out of the company, and attack other companies.
Of even greater concern - if the trend continues for another year, we will see sometime in 2027 AI systems that are powerful enough to go on very long-range attacks with very top-tier cyber capabilities in a way that is going to be much more difficult to detect and shut off. We're still struggling to contain today's systems, let alone tomorrow's.